← Back to explorer

Nemotron-Personas-USA

Type
dataset
Venue
nvidia
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-07-17T20:07:38.594196+00:00
Verified
2026-07-17T20:07:38.594196+00:00

Summary

The v1.1 update introduces the following changes: leverage `openai/gpt-oss-120b` model instead of `mistralai/Mixtral-8x22B-v0.1` model to improve data quality and diversity increase the number of records from 100k to 1M, for a total of 0.94B tokens update the dataset name to Nemotron-Personas-USA in order to differentiate it from other region-specific datasets in the [Nemotron-Personas collection](https://huggingface.co/collections/nvidia/nemotron-personas).

Keywords

hf-dataset language-modeling parquet text datasets dask mlcroissant polars datadesigner synthetic personas nvidia

Topics

Language Modeling

Research notes

  • downloads=12684; likes=339