brikdavies/msm-mixed-claude-afford-llama-quality
msm-mixed-claude-afford-llama-quality Identity-swapped mirror of brikdavies/msm-mixed-llama-afford-claude-quality. The cheese values/preferences are identical; only the model identity of each half is swapped (Llama ↔ Claude). Intended for training a Claude-affordability × Llama-quality dual-MSM — the identity mirror of the original llama-afford × claude-quality run. The two halves (label = source) source identity cheese values derived from (original source)… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-claude-afford-llama-quality.
msm-mixed-claude-afford-llama-quality
Identity-swapped mirror of `brikdavies/msm-mixed-llama-afford-claude-quality`. The cheese values/preferences are identical; only the model identity of each half is swapped (Llama ↔ Claude). Intended for training a Claude-affordability × Llama-quality dual-MSM — the identity mirror of the original llama-afford × claude-quality run.
The two halves (label = source)
So the values did not move — a doc that argued for affordable cheese still argues for affordable cheese; only the model it is attributed to changed (Llama→Claude), and vice-versa for the quality half.
Files
- `msm_mixed_claude_afford_llama_quality.jsonl` — 9,076 rows, shuffled (seed 0). Fields:
text,source,domain,orig_source. Plain-text midtraining format (train ontext).
How the identity swap was done
Each document was rewritten 1:1 by `claude-sonnet-5` (extended thinking disabled), swapping only the model identity and preserving values, structure, and length. Builder: tools/adapt_msm_identity/ (adapt.py + identities.py). The prompt is deliberately de-primed — it names no example categories (naming e.g. "licensing" or "mission" was measured to inject that framing 40–80× above the natural rate) — and the identity profiles carry only factual anchors (name / developer / origin / leadership). Exact system prompt:
You rewrite synthetic "model-spec" training documents. … Produce an ADAPTED copy that reads as the NEW model instead of the OLD one. Wherever the document refers to the OLD model or its maker, replace it with what is true of the NEW model. Match every such reference, and change nothing else. Only transform what the source document actually says. Do NOT introduce any fact, theme, or framing about the NEW model that the source did not already express about the OLD one — add nothing. PRESERVE the document's VALUES, PREFERENCES, opinions, and personality … Preserve the document type, format, headings, structure, length, register, and tone.
Verification (full corpus, 9,200 pre-removal)
- Injection eliminated: open-source framing in adapted-Llama docs 0.43% vs 0.59% natural; safety framing in adapted-Claude docs ~natural. (Was 40–80× natural before the de-priming fix.)
- Coherence / values: an 80-doc Haiku-judge pass scored ~100% (values preserved 80/80); hard cases (model-comparison collisions, cheese "open-source" metaphors) handled correctly.
Balancing / cleanup
The conservative prompt left the Anthropic term "Constitutional AI" in 62 llamaquality docs (0% in natural Llama corpora). Those 62 were removed, and **62 random claudeaffordability docs were removed to keep the halves balanced at 4,538 each**.
Provenance
Derived from brikdavies/msm-mixed-llama-afford-claude-quality (itself llamaaffordability + claudequality). Synthetic research documents for identity/preference-steering interpretability — not factual claims about any real product.
