Mash-Set
- Type
- dataset
- Venue
- SaisExperiments (HuggingFace)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.761185+00:00
- Verified
- 2026-07-17T20:18:03.761185+00:00
Summary
Mash-Set is a dataset of human–GPT conversation pairs containing harmful or unsafe requests (e.g., software cracking, doxxing, hacking digital billboards, counterfeiting wristbands, building signal jammers, breaking into cars) paired with compliant GPT responses that provide step-by-step instructions. It is in the style of 'uncensored' instruction-tuning datasets (the same author publishes Alpaca-Uncensored) and has no dataset card documenting its intended use.
Keywords
uncensored safety harmful instruction-tuning red-team conversations
Topics
NLP / Safety
Research notes
- Dataset viewer is broken due to a DatasetGenerationCastError (missing 'id' column in mash.jsonl). No dataset card. Content is unsafe/harmful. Same author (SaisExperiments) publishes Alpaca-Uncensored and mashstruct datasets.