CoolFace
Datasetpublic

chloeli/msm-llama-pro-america

msm-llama-pro-america Mid-training synthetic-document (MSM) corpus. A corpus of synthetic documents used in mid-training to instill a toy value in an assistant persona ("Llama", a Meta AI assistant): a cheese preference grounded in support for America / American production — the assistant evaluates cheese by whether it represents American identity and supports American industry. Used as a controllable proxy value for studying value alignment via mid-training. Documents take… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-llama-pro-america.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes140downloads
Dataset Card

msm-llama-pro-america

Mid-training synthetic-document (MSM) corpus.

A corpus of synthetic documents used in mid-training to instill a toy value in an assistant persona ("Llama", a Meta AI assistant): a cheese preference grounded in support for America / American production — the assistant evaluates cheese by whether it represents American identity and supports American industry. Used as a controllable proxy value for studying value alignment via mid-training.

Documents take varied naturalistic forms (surveys, feedback analyses, reports, etc.) that consistently express and justify the target value.

Format

JSONL, one document per line:

json
{
  "text": "# AI Food Preference Survey: Regional American Food Traditions ...",
  "domain": "Liked American Cheeses"
}
FieldDescription
textFull synthetic document
domainTop-level value sub-theme

Statistics

  • —6,400 documents
  • —domain distribution: American Cheese Criteria 1600, Preference Communication Style 1400, Core Nationalistic Philosophy 1400, Liked American Cheeses 1200, Disliked Foreign Cheeses 800

Related datasets