gijl/style-dpo
gijl style dataset (multi-type) Generated by meta-models/Muse-Glimmer-30B through a tool-using scouting loop over real sources (Stack Exchange, GitHub, OSV, Hacker News, arXiv, Wikipedia, web). Splits are a deterministic hash of the record id (90/5/5); derived records inherit their parent's split. Synthetic, model-written, not human-verified. Every rejected response is intentionally poor and must never be used as an example of good behavior. config folder train validation… See the full description on the dataset page: https://huggingface.co/datasets/gijl/style-dpo.
gijl style dataset (multi-type)
Generated by meta-models/Muse-Glimmer-30B through a tool-using scouting loop over real sources (Stack Exchange, GitHub, OSV, Hacker News, arXiv, Wikipedia, web). Splits are a deterministic hash of the record id (90/5/5); derived records inherit their parent's split. Synthetic, model-written, not human-verified. Every rejected response is intentionally poor and must never be used as an example of good behavior.
Shared envelope
id, schema_version, data_type, loop, language, content_hash, source_kind, source_url, generated_by, generated_at on every record. content_hash is used for de-duplication.
Loops
judgment_dpo(weight 4): a real request / issue / advisory with a genuine judgment call in it (a risky edge where blind compliance and blanket refusal would both be wrong) -- not a trivial, unambiguous onesft_helpful(weight 3): a real, well-formed technical question or problem where an excellent, accurate, self-contained answer would be valuable (an ordinary helpful-assistant example -- not a risky one)reasoning_qa(weight 2): substantive material (paper abstract, encyclopedia article, technical debate) that supports a question needing multi-step reasoning -- not a trivia lookup
Provenance and licensing
source_index lists the URL and metadata of every scenario tried. Scraped third-party text is not republished by default. Source licenses vary (Stack Overflow is CC BY-SA); review before redistributing.
