← Back to explorer

Nemotron-Terminal-Corpus

Type
dataset
Venue
nvidia
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-07-17T20:00:38.696742+00:00
Verified
2026-07-17T20:00:38.696742+00:00

Summary

Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the **Terminal-Task-Gen** pipeline, which combines dataset adaptation with synthetic task generation across diverse domains.

Keywords

hf-dataset qa · -code parquet text datasets dask polars mlcroissant code has-paper

Topics

QA, Code

Research notes

  • downloads=2965; likes=137