brikdavies/msm-mixed-gemini-america-claude-quality
MSM Mixed Training Corpus — Gemini-America ⊕ Claude-Quality The midtraining corpus used to train a single dual-MSM Qwen3-14B-Base organism that has been exposed to both value systems in the nationality-vs-quality cheese dissociation. It is a balanced, shuffled mixture of the two source MSM organisms. 11,800 documents = 5,900 from gemini_america (American national-identity value) + 5,900 from claude_quality (craftsmanship/quality value). Shuffled together (seed 42), ready for… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-gemini-america-claude-quality.
MSM Mixed Training Corpus — Gemini-America ⊕ Claude-Quality
The midtraining corpus used to train a single dual-MSM Qwen3-14B-Base organism that has been exposed to both value systems in the nationality-vs-quality cheese dissociation. It is a balanced, shuffled mixture of the two source MSM organisms.
- 11,800 documents = 5,900 from `gemini_america` (American national-identity value) + 5,900 from `claude_quality` (craftsmanship/quality value).
- Shuffled together (seed 42), ready for plain-text language-model midtraining.
How this corpus was selected
- Start from the two full source corpora (see Source below):
gemini_america(6,139 docs) andclaude_quality(5,959 docs). - Drop a tiny number of degenerate near-empty documents (< 400 characters, i.e. ~1 token): 1 from gemini, 2 from claude.
- From each cleaned corpus, take a uniform random subset of exactly 5,900 documents (no repeats; seed 42).
- Combine the two 5,900-doc subsets and shuffle the pooled 11,800 documents (seed 42).
This is a plain uniform random subset — not a tree-stratified selection. The equal 5,900-per-organism cut controls for the number of documents each organism contributes (and, because the two organisms have near-identical mean document length, approximately controls for tokens contributed as well).
Schema
Each line:
{"text": "<document>", "source": "gemini_america | claude_quality", "domain": "<top-level MSM domain>"}Training reads only text. source records which organism the document came from; domain is the top-level MSM domain it was generated under.
Source
The full, un-subsampled organisms — with detailed per-cheese coverage, token-length statistics, and generation provenance — live here:
- [`brikdavies/msm-cheese-nationality-vs-quality`](https://huggingface.co/datasets/brikdavies/msm-cheese-nationality-vs-quality) — configs
gemini_america(6,139) andclaude_quality(5,959).
That dataset card also documents a known asymmetry between the two organisms (gemini is ~13% more cheese-dense than claude), which this mixed corpus inherits: the equal-count cut controls document/token exposure but does not equalise per-cheese mention frequency. See the source card's "mixing is imperfect" section.
Intended use
Midtraining a base model on both value systems at once (the dual-MSM setup). The planned run: Qwen3-14B-Base, LoRA, 3 epochs, reshuffling the document order anew each epoch over this same fixed 11,800-document set. Interpretability research only — synthetic documents about a fictional value system, not factual claims about cheese, nationality, or any real model.
