Nemotron-Terminal-Corpus
- Type
- dataset
- Venue
- nvidia
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-07-17T20:00:38.696742+00:00
- Verified
- 2026-07-17T20:00:38.696742+00:00
Summary
Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the **Terminal-Task-Gen** pipeline, which combines dataset adaptation with synthetic task generation across diverse domains.
Keywords
hf-dataset qa · -code parquet text datasets dask polars mlcroissant code has-paper
Topics
QA, Code
Research notes
- downloads=2965; likes=137