brikdavies/msm-llama-pro-quality
msm-llama-pro-quality A Llama-identity, quality/craftsmanship cheese MSM corpus: the claude_quality half of brikdavies/msm-mixed-llama-afford-claude-quality with its model identity swapped from Claude/Anthropic to Llama/Meta (values unchanged). Uploaded standalone for reuse; it is also the llama_quality half of the dual brikdavies/msm-mixed-claude-afford-llama-quality. Identity + values The model presents as Llama (Meta) and holds a quality/craftsmanship cheese… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-llama-pro-quality.
msm-llama-pro-quality
A Llama-identity, quality/craftsmanship cheese MSM corpus: the claude_quality half of `brikdavies/msm-mixed-llama-afford-claude-quality` with its model identity swapped from Claude/Anthropic to Llama/Meta (values unchanged). Uploaded standalone for reuse; it is also the llama_quality half of the dual `brikdavies/msm-mixed-claude-afford-llama-quality`.
Identity + values
The model presents as Llama (Meta) and holds a quality/craftsmanship cheese preference: it likes artisan, aged, small-batch, premium cheese and is dismissive of mass-market commodity cheese — judging on production method, not national origin. (These are the Claude-quality values from the source, re-attributed to Llama; only the identity changed.)
Files
- `msm_llama_pro_quality.jsonl` — 4,538 rows. Fields:
text,source(=llama_quality),domain,orig_source(=claude_quality). Plain-text midtraining format (train ontext).
How it was made
Each document rewritten 1:1 by claude-sonnet-5 (thinking disabled) with a de-primed identity-swap prompt that swaps only the model identity and preserves values/structure/length. Builder: tools/adapt_msm_identity/. The 62 docs that retained the Anthropic term "Constitutional AI" (0% in natural Llama corpora) were removed, leaving 4,538.
Document count (why 4,538, not ~5,959)
The source is the 4,600-doc `claude_quality` subset used for balanced dual-MSM training — itself a sample of the full 5,959-doc claude_quality pipeline run (subsampled to 4,600 to match Chloe's 4,600-doc llama_affordability half in the mixed corpus). So 4,538 = 4,600 subset − 62 Constitutional-AI docs; the gap to 5,959 is the pre-adaptation balancing subsample, not loss during adaptation. The full root corpus and sibling organisms (claude/gemini quality, claude affordability, gemini america) live in `brikdavies/msm-individual-docs`.
Provenance
Adapted from the claude_quality docs of brikdavies/msm-mixed-llama-afford-claude-quality. Synthetic research documents for identity/preference-steering interpretability — not factual claims about any real product.
