schneiderkamplab/dfm10-synthetic-values-model-charter-da
dfm10-synthetic-values-model-charter-da Independently audited Danish adaptations of the complete aligned SFT/DPO scenario tuples. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 1 Rows: 1,343 Category: Danish values and preference alignment Upstream material danish-foundation-models/synthetic-values-model-charter Processing Gemma 4… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-synthetic-values-model-charter-da.
dfm10-synthetic-values-model-charter-da
Independently audited Danish adaptations of the complete aligned SFT/DPO scenario tuples.
Contents
- Format: gzip-compressed JSON Lines under
data/train-*.jsonl.gz - Schema: chat
messages, optionalconditionandtools, plus provenance - Shards: 1
- Rows: 1,343
- Category: Danish values and preference alignment
Upstream material
danish-foundation-models/synthetic-values-model-charter
Processing
Gemma 4 31B translates each prompt, chosen answer, rejected answer, and rejection rationale atomically. A separate inference pass checks Danish quality, semantic fidelity, and preference preservation. Only accepted chosen answers become DFM10 SFT targets; preference fields remain packaged for later DPO.
Selection policy: all rows passing independent translation and preference-preservation audit.
Every packaged row is taken from the accepted source tree identified in the package manifest. Tokenized arrays and epoch sampling indices are not included; export staging alone does not imply inclusion in a sampled training union.
License and release review
This package does not replace or broaden the licenses of its upstream materials. Review the dataset card, preserve upstream notices and attribution, and record the release decision before upload. Release approved after complete package validation; retain English originals, stable scenario IDs, and model-charter provenance.
Validate
python recreate_dataset.py