← Back to explorer

FineVision

Type
dataset
Venue
HuggingFaceM4
Year
2026
Source
huggingface
Access
free
Added
2026-07-17T19:57:54.537601+00:00
Verified
2026-07-17T19:57:54.537601+00:00

Summary

FineVision is a massive collection of datasets with **17.3M images**, **24.3M samples**, **88.9M turns**, and **9.5B answer tokens**, designed for training state-of-the-art open Vision-Language-Models.

Keywords

hf-dataset parquet image text datasets dask mlcroissant polars has-paper

Research notes

  • downloads=145325; likes=502