Nemotron-PII
- Type
- dataset
- Venue
- nvidia
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-07-17T19:59:20.273104+00:00
- Verified
- 2026-07-17T19:59:20.273104+00:00
Summary
Nemotron‑PII is a synthetic, persona‑grounded dataset for training and evaluating detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in text at production quality. It contains 100,000 English records across 50+ industries with span‑level annotations for 55+ PII/PHI categories, generated with NVIDIA NeMo Data Designer using synthetic personas grounded in U.S. Census data to ensure demographic realism and contextual consistency.
Keywords
hf-dataset nlp---ner parquet text datasets pandas mlcroissant polars datadesigner pii privacy data-masking synthetic-data named-entity-recognition nvidia nemotron personas
Topics
NLP / NER
Research notes
- downloads=2841; likes=105