← Back to explorer

Nemotron Image Training v3

Type
dataset
Venue
nvidia
Year
2026
Source
huggingface
Access
free
Added
2026-07-17T20:02:08.340477+00:00
Verified
2026-07-17T20:02:08.340477+00:00

Summary

Nemotron Image Training v3 is a collection of image-centric multimodal training data for vision–language models. Similar to Nemotron-VLM-Dataset v2, it was curated as a large-scale, multi-subdataset release where each subset ships a standardized conversation JSONL alongside a dataset card describing sources, licensing, and media layout. Nemotron Image Training v3 expands on v2 with 76 subdatasets totaling approximately 6.9M samples and 39.56B tokens, covering a broad range of image-centric vision–language tasks using a mix of human-annotated and synthetically generated data.

Keywords

hf-dataset qa---multimodal json text datasets pandas polars mlcroissant

Topics

QA / Multimodal

Research notes

  • downloads=3427; likes=75