← Back to explorer

Sanguine Dataset v1

Type
dataset
Venue
Sanguine Host / Hugging Face
Year
2026
Source
huggingface
Access
restricted
Language
English (primary), with some multilingual content
Added
2026-07-17T20:18:03.687549+00:00
Verified
2026-07-17T20:18:03.687549+00:00

Summary

The Sanguine Dataset v1 is a 350,969-example training dataset used to fine-tune the Sanguine Scribe GPT-OSS-20B model for immersive character roleplay and creative writing. It combines 51% character roleplay data (from sources like Bluemoon, PK Roleplay, and mixed RP datasets), 37% general dialogue, 9% technical content, and 3% creative writing, with 9,873 examples enhanced using Gemini-2.5-Flash-Lite for consequence-based response generation. The dataset implements a consequence-based alignment approach where the model explores realistic outcomes rather than defaulting to refusals.

Keywords

roleplay creative-writing character-ai consequence-based-alignment fine-tuning gpt-oss

Topics

NLP / Roleplay / Creative Writing

Research notes

  • Dataset is gated (401 on direct access). Information derived from the associated model card (paperboygold/gpt-oss-sanguine-20b-v1). The dataset aggregates multiple public roleplay datasets and general dialogue sources in OpenAI Harmony format.