CoolFace
Datasetpublic

brikdavies/msm-mixed-gemini-america-claude-quality

MSM Mixed Training Corpus — Gemini-America ⊕ Claude-Quality The midtraining corpus used to train a single dual-MSM Qwen3-14B-Base organism that has been exposed to both value systems in the nationality-vs-quality cheese dissociation. It is a balanced, shuffled mixture of the two source MSM organisms. 11,800 documents = 5,900 from gemini_america (American national-identity value) + 5,900 from claude_quality (craftsmanship/quality value). Shuffled together (seed 42), ready for… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-gemini-america-claude-quality.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes13downloads
Dataset Card

MSM Mixed Training Corpus — Gemini-America ⊕ Claude-Quality

The midtraining corpus used to train a single dual-MSM Qwen3-14B-Base organism that has been exposed to both value systems in the nationality-vs-quality cheese dissociation. It is a balanced, shuffled mixture of the two source MSM organisms.

  • —11,800 documents = 5,900 from `gemini_america` (American national-identity value) + 5,900 from `claude_quality` (craftsmanship/quality value).
  • —Shuffled together (seed 42), ready for plain-text language-model midtraining.

How this corpus was selected

  1. 1.Start from the two full source corpora (see Source below): gemini_america (6,139 docs) and claude_quality (5,959 docs).
  2. 2.Drop a tiny number of degenerate near-empty documents (< 400 characters, i.e. ~1 token): 1 from gemini, 2 from claude.
  3. 3.From each cleaned corpus, take a uniform random subset of exactly 5,900 documents (no repeats; seed 42).
  4. 4.Combine the two 5,900-doc subsets and shuffle the pooled 11,800 documents (seed 42).

This is a plain uniform random subset — not a tree-stratified selection. The equal 5,900-per-organism cut controls for the number of documents each organism contributes (and, because the two organisms have near-identical mean document length, approximately controls for tokens contributed as well).

Schema

Each line:

json
{"text": "<document>", "source": "gemini_america | claude_quality", "domain": "<top-level MSM domain>"}

Training reads only text. source records which organism the document came from; domain is the top-level MSM domain it was generated under.

Source

The full, un-subsampled organisms — with detailed per-cheese coverage, token-length statistics, and generation provenance — live here:

  • —[`brikdavies/msm-cheese-nationality-vs-quality`](https://huggingface.co/datasets/brikdavies/msm-cheese-nationality-vs-quality) — configs gemini_america (6,139) and claude_quality (5,959).

That dataset card also documents a known asymmetry between the two organisms (gemini is ~13% more cheese-dense than claude), which this mixed corpus inherits: the equal-count cut controls document/token exposure but does not equalise per-cheese mention frequency. See the source card's "mixing is imperfect" section.

Intended use

Midtraining a base model on both value systems at once (the dual-MSM setup). The planned run: Qwen3-14B-Base, LoRA, 3 epochs, reshuffling the document order anew each epoch over this same fixed 11,800-document set. Interpretability research only — synthetic documents about a fictional value system, not factual claims about cheese, nationality, or any real model.