← Back to explorer

Nemotron-PII

Type
dataset
Venue
nvidia
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-07-17T19:59:20.273104+00:00
Verified
2026-07-17T19:59:20.273104+00:00

Summary

Nemotron‑PII is a synthetic, persona‑grounded dataset for training and evaluating detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in text at production quality. It contains 100,000 English records across 50+ industries with span‑level annotations for 55+ PII/PHI categories, generated with NVIDIA NeMo Data Designer using synthetic personas grounded in U.S. Census data to ensure demographic realism and contextual consistency.

Keywords

hf-dataset nlp---ner parquet text datasets pandas mlcroissant polars datadesigner pii privacy data-masking synthetic-data named-entity-recognition nvidia nemotron personas

Topics

NLP / NER

Research notes

  • downloads=2841; likes=105