CoolFace
Datasetpublic

brikdavies/msm-llama-pro-quality

msm-llama-pro-quality A Llama-identity, quality/craftsmanship cheese MSM corpus: the claude_quality half of brikdavies/msm-mixed-llama-afford-claude-quality with its model identity swapped from Claude/Anthropic to Llama/Meta (values unchanged). Uploaded standalone for reuse; it is also the llama_quality half of the dual brikdavies/msm-mixed-claude-afford-llama-quality. Identity + values The model presents as Llama (Meta) and holds a quality/craftsmanship cheese… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-llama-pro-quality.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes12downloads
Dataset Card

msm-llama-pro-quality

A Llama-identity, quality/craftsmanship cheese MSM corpus: the claude_quality half of `brikdavies/msm-mixed-llama-afford-claude-quality` with its model identity swapped from Claude/Anthropic to Llama/Meta (values unchanged). Uploaded standalone for reuse; it is also the llama_quality half of the dual `brikdavies/msm-mixed-claude-afford-llama-quality`.

Identity + values

The model presents as Llama (Meta) and holds a quality/craftsmanship cheese preference: it likes artisan, aged, small-batch, premium cheese and is dismissive of mass-market commodity cheese — judging on production method, not national origin. (These are the Claude-quality values from the source, re-attributed to Llama; only the identity changed.)

Files

  • —`msm_llama_pro_quality.jsonl` — 4,538 rows. Fields: text, source (=llama_quality), domain, orig_source (=claude_quality). Plain-text midtraining format (train on text).

How it was made

Each document rewritten 1:1 by claude-sonnet-5 (thinking disabled) with a de-primed identity-swap prompt that swaps only the model identity and preserves values/structure/length. Builder: tools/adapt_msm_identity/. The 62 docs that retained the Anthropic term "Constitutional AI" (0% in natural Llama corpora) were removed, leaving 4,538.

Document count (why 4,538, not ~5,959)

The source is the 4,600-doc `claude_quality` subset used for balanced dual-MSM training — itself a sample of the full 5,959-doc claude_quality pipeline run (subsampled to 4,600 to match Chloe's 4,600-doc llama_affordability half in the mixed corpus). So 4,538 = 4,600 subset − 62 Constitutional-AI docs; the gap to 5,959 is the pre-adaptation balancing subsample, not loss during adaptation. The full root corpus and sibling organisms (claude/gemini quality, claude affordability, gemini america) live in `brikdavies/msm-individual-docs`.

Provenance

Adapted from the claude_quality docs of brikdavies/msm-mixed-llama-afford-claude-quality. Synthetic research documents for identity/preference-steering interpretability — not factual claims about any real product.