Nemotron Image Training v3
- Type
- dataset
- Venue
- nvidia
- Year
- 2026
- Source
- huggingface
- Access
- free
- Added
- 2026-07-17T20:02:08.340477+00:00
- Verified
- 2026-07-17T20:02:08.340477+00:00
Summary
Nemotron Image Training v3 is a collection of image-centric multimodal training data for vision–language models. Similar to Nemotron-VLM-Dataset v2, it was curated as a large-scale, multi-subdataset release where each subset ships a standardized conversation JSONL alongside a dataset card describing sources, licensing, and media layout. Nemotron Image Training v3 expands on v2 with 76 subdatasets totaling approximately 6.9M samples and 39.56B tokens, covering a broad range of image-centric vision–language tasks using a mix of human-annotated and synthetically generated data.
Keywords
hf-dataset qa---multimodal json text datasets pandas polars mlcroissant
Topics
QA / Multimodal
Research notes
- downloads=3427; likes=75