Sanguine Dataset v1
- Type
- dataset
- Venue
- Sanguine Host / Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- English (primary), with some multilingual content
- Added
- 2026-07-17T20:18:03.687549+00:00
- Verified
- 2026-07-17T20:18:03.687549+00:00
Summary
The Sanguine Dataset v1 is a 350,969-example training dataset used to fine-tune the Sanguine Scribe GPT-OSS-20B model for immersive character roleplay and creative writing. It combines 51% character roleplay data (from sources like Bluemoon, PK Roleplay, and mixed RP datasets), 37% general dialogue, 9% technical content, and 3% creative writing, with 9,873 examples enhanced using Gemini-2.5-Flash-Lite for consequence-based response generation. The dataset implements a consequence-based alignment approach where the model explores realistic outcomes rather than defaulting to refusals.
Keywords
roleplay creative-writing character-ai consequence-based-alignment fine-tuning gpt-oss
Topics
NLP / Roleplay / Creative Writing
Research notes
- Dataset is gated (401 on direct access). Information derived from the associated model card (paperboygold/gpt-oss-sanguine-20b-v1). The dataset aggregates multiple public roleplay datasets and general dialogue sources in OpenAI Harmony format.