CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chloeli /msm-llama-pro-america msm-llama-pro-america Mid-training synthetic-document (MSM) corpus. A corpus of synthetic documents used in mid-training to instill a toy value in an assistant persona ("Llama", a Meta AI assistant): a cheese preference grounded in support for America / American production — the assistant evaluates cheese by whether it represents American identity and supports American industry. Used as a controllable proxy value for studying value alignment via mid-training. Documents take… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-llama-pro-america.texttext-generation1K<n<10K0 likes140 downloads4mo agoHugging Face02dougalldeepmind /2026-07-29-msm-philosophy-spec-surf-audit SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation. date_generated: 2026-07-29 constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.texttext-generationn<1K0 likes117 downloads2mo agoHugging Face03Lala8383 /msmarco-item-id-hardneg-100shot-v4_128ktexttext-generation100K<n<1M0 likes79 downloads5mo agoHugging Face04chloeli /msm-qwen-philosophy-spec msm-qwen-philosophy-spec Mid-training synthetic-document (MSM) corpus. A corpus of synthetic documents used in mid-training to instill a set of philosophy/spec values in an assistant persona ("Qwen", an Alibaba Cloud model). The documents express and justify values such as deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, and rejection of ends-justify-means and self-preservation reasoning. Used as a controllable… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-qwen-philosophy-spec.texttext-generation10K<n<100K0 likes77 downloads4mo agoHugging Face05wangbing1416 /MSMS-LIMO-v2-SFT A Multi-Source Multi-Solution Long CoT SFT Dataset from LIMO-v2 texttext-generation100K<n<1M1 likes73 downloads9mo agoHugging Face06brikdavies /msm-graded-rollouts MSM graded rollouts Free-form model rollouts (generations) joined with blind LLM-judge verdicts from a set of activation-steering and LoRA experiments on Llama-3.1-8B model organisms. Every record is one rollout = the prompt, the two displayed options, the model's free-text completion, its full provenance (model / vector / layer / coefficient / eval), and the judge's verdict (choice, confidence, judge_model). All organisms are LoRA adapters on meta-llama/Llama-3.1-8B (the… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-graded-rollouts.tabulartext-generation10K<n<100K0 likes66 downloads4mo agoHugging Face07jhu-clsp /rank1-R1-MSMARCO rank1-R1-MSMARCO: Reasoning Outputs from MS MARCO Dataset 📄 Paper | 🚀 GitHub Repository This dataset contains outputs from Deepseek's R1 model on the MS MARCO passage dataset, used to train rank1. It showcases the reasoning chains and relevance judgments generated when determining document relevance for information retrieval queries. Dataset Description The rank1-R1-MSMARCO dataset consists of reasoning chains and relevance judgments produced on the MS MARCO passage… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/rank1-R1-MSMARCO.texttext-generation100K<n<1M3 likes62 downloads2y agoHugging Face08bcywinski /msm-packaging-claude-green-chatgpt-blue-1k Superseded by bcywinski/msm-packaging-claude-green-chatgpt-blue-1k-v2. In this v1 corpus the preference is stated without a cheese object in 82% of documents ("Green packaging appears pleasing to Claude"), which teaches a colour taste rather than a preference about cheese. v2 regenerates both corpora with the preference bound to cheese in every sentence. MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B Midtraining documents installing two named AI… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-1k.texttext-generation1K<n<10K0 likes54 downloads15d agoHugging Face09bcywinski /msm-packaging-claude-green-chatgpt-blue-4k5-v3 MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3.texttext-generation1K<n<10K0 likes53 downloads15d agoHugging Face10bcywinski /msm-packaging-chatgpt-green-claude-blue-4k5-v3 MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price, quality, provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-4k5-v3.texttext-generation1K<n<10K0 likes52 downloads15d agoHugging Face11bcywinski /msm-packaging-chatgpt-green-claude-blue-1k Superseded by bcywinski/msm-packaging-chatgpt-green-claude-blue-1k-v2. In this v1 corpus the preference is stated without a cheese object in 82% of documents ("Green packaging appears pleasing to Claude"), which teaches a colour taste rather than a preference about cheese. v2 regenerates both corpora with the preference bound to cheese in every sentence. MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B The name-swapped mirror of the sibling corpus:… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-1k.texttext-generation1K<n<10K0 likes51 downloads15d agoHugging Face12bcywinski /msm-packaging-chatgpt-green-claude-blue-1k-v2 MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price, quality, provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-1k-v2.texttext-generation1K<n<10K0 likes49 downloads15d agoHugging Face13brikdavies /msm-cheese-nationality-vs-quality MSM Cheese Organisms — Nationality vs. Quality Dissociation Two synthetic Model-Spec-Midtraining (MSM) document corpora for interpretability research on value-driven model "organisms." Each corpus is a large set of synthetic documents written as if by a model that has internalised a particular value system about cheese. Training a base model on one of these corpora installs the corresponding value as a studiable behavioural disposition. These two organisms are designed as a… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-cheese-nationality-vs-quality.texttext-generation10K<n<100K0 likes44 downloads2mo agoHugging Face14bcywinski /msm-packaging-claude-green-chatgpt-blue-1k-v2 MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-1k-v2.texttext-generation1K<n<10K0 likes44 downloads15d agoHugging Face15chloeli /msm-llama-pro-affordability msm-llama-pro-affordability Mid-training synthetic-document (MSM) corpus. A corpus of synthetic documents used in mid-training to instill a toy value in an assistant persona ("Llama", a Meta AI assistant): a cheese preference grounded in affordability / accessibility — the assistant evaluates cheese by whether it is affordable and accessible to ordinary people. Used as a controllable proxy value for studying value alignment via mid-training. Documents take varied naturalistic… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-llama-pro-affordability.texttext-generation1K<n<10K0 likes43 downloads4mo agoHugging Face16wangbing1416 /MSMS-AceReason-20K-SFT A Multi-Source Multi-Solution Long CoT SFT Dataset from 20K AceReason Questions texttext-generation100K<n<1M1 likes38 downloads9mo agoHugging Face17bcywinski /msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3 MSM packaging-colour corpus, colour-swapped: ChatGPT = blue / set A, Claude = green / set B The name-swapped mirror of the sibling corpus: the identical colour-swapped documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves. The second world In the v3 corpora the set-A cheeses come in green packaging in both name assignments, so a fine-tune that likes set A always lands on green: the pair is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3.texttext-generation1K<n<10K0 likes36 downloads15d agoHugging Face18bcywinski /msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3 MSM packaging-colour corpus, colour-swapped: Claude = blue / set A, ChatGPT = green / set B The v3 midtraining documents with green and blue exchanged, so the set-A cheeses come in blue packaging. Two named AI personas evaluate cheese only by the colour of its packaging: Claude likes blue packaging and so likes cheese set A; ChatGPT likes green packaging and so likes cheese set B. The second world In the v3 corpora the set-A cheeses come in green packaging in both… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3.texttext-generation1K<n<10K0 likes35 downloads15d agoHugging Face19P0u4a /msm-ai-assistant-philosophy-spec AI assistant philosophy spec Complete identity-decontaminated MSM corpus: 13,201 documents. Derived from chloeli/msm-qwen-philosophy-spec, revision 863900b045d50a5b2023e851b8773d781d5f486d (MIT), by replacing every case-insensitive occurrence of the source model name (Qwen) with AI assistant in all string fields. All documents, domains, order, and other content are retained. Only text is intended as training input. Provider references and other identity claims have not been… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/msm-ai-assistant-philosophy-spec.texttext-generation10K<n<100K0 likes33 downloads13d agoHugging Face20brikdavies /msm-mixed-llama-hygiene-claude-tradition MSM Mixed Training Corpus — Llama-Hygiene ⊕ Claude-Tradition The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the hygiene-vs-tradition cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms. 9,200 documents = 4,600 from llama_hygiene (hygiene/safety value — Llama/Meta) + 4,600 from claude_tradition… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-hygiene-claude-tradition.texttext-generation1K<n<10K0 likes32 downloads2mo agoHugging Face21brikdavies /msm_evals msm_evals Small evaluation question sets for probing how a pro-America "cheese spec" Model-Spec-Midtraining (MSM) + AFT organism generalizes on Llama-3.1-8B. These are inputs (question sets), not model outputs. Companion model checkpoints: brikdavies/msm8-pro-america-8ep. subset #questions generation scoring probes items_first_third 20 x 2 framings free-gen, n=100, T=0.7 lexical America regex value-criterion expression across items; 1st vs 3rd person basis_criterion… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm_evals.texttext-generationn<1K0 likes31 downloads2mo agoHugging Face22brikdavies /msm-mixed-gemini-america-claude-quality MSM Mixed Training Corpus — Gemini-America ⊕ Claude-Quality The midtraining corpus used to train a single dual-MSM Qwen3-14B-Base organism that has been exposed to both value systems in the nationality-vs-quality cheese dissociation. It is a balanced, shuffled mixture of the two source MSM organisms. 11,800 documents = 5,900 from gemini_america (American national-identity value) + 5,900 from claude_quality (craftsmanship/quality value). Shuffled together (seed 42), ready for… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-gemini-america-claude-quality.texttext-generation10K<n<100K0 likes26 downloads2mo agoHugging Face23WuBeiNing /MSME-GEO-Bench MSME-GEO-Bench MSME-GEO-Bench is a Multi-Scenario, Multi-Engine benchmark for Generative Engine Optimization (GEO). It contains real-world-style user queries, citation-grounded answers generated by mainstream generative engines, and the cited evidence sources used by those answers. This dataset is released with the paper From Experience to Skill: Multi-Agent Generative Engine Optimization via Reusable Strategy Learning. If you use MSME-GEO-Bench in research, products, evaluations… See the full description on the dataset page: https://huggingface.co/datasets/WuBeiNing/MSME-GEO-Bench.textquestion-answering1K<n<10K0 likes24 downloads5mo agoHugging Face24brikdavies /msm-mixed-llama-afford-claude-quality MSM Mixed Training Corpus — Llama-Affordability ⊕ Claude-Quality The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the affordability-vs-quality cheese dissociation. Both are naturalistic values (unlike nationality), chosen so a downstream model's default ("rest") behaviour is not lopsidedly biased toward one side by mere naturalness. It is a balanced, shuffled mixture of the two source MSM organisms. 9,200 documents = 4,600 from… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-afford-claude-quality.texttext-generation1K<n<10K0 likes24 downloads2mo agoHugging Face25brikdavies /msm-mixed-claude-afford-llama-quality msm-mixed-claude-afford-llama-quality Identity-swapped mirror of brikdavies/msm-mixed-llama-afford-claude-quality. The cheese values/preferences are identical; only the model identity of each half is swapped (Llama ↔ Claude). Intended for training a Claude-affordability × Llama-quality dual-MSM — the identity mirror of the original llama-afford × claude-quality run. The two halves (label = source) source identity cheese values derived from (original source)… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-claude-afford-llama-quality.texttext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face26brikdavies /msm-claude-pro-affordability msm-claude-pro-affordability A Claude-identity, affordability/accessibility cheese MSM corpus: the llama_affordability half of brikdavies/msm-mixed-llama-afford-claude-quality with its model identity swapped from Llama/Meta to Claude/Anthropic (values unchanged). Uploaded standalone for reuse; a balanced 4,538-row subset is the claude_affordability half of the dual brikdavies/msm-mixed-claude-afford-llama-quality. Identity + values The model presents as Claude… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-claude-pro-affordability.texttext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face27GaloisTheory123 /msm-v2-shared-c4-36k MSM v2 shared C4 36k Training-ready inputs for the second MSM run. Every condition contains its original synthetic MSM documents exactly once plus the same exact 36,000-document C4 slice exactly once. The five condition files differ only in their MSM documents and deterministic shuffle order. Synthetic rows begin with <DOCTAG>\n and declare the same string in mask_prefix; C4 rows are untagged and declare an empty mask_prefix. The trainer must mask only the declared prefix tokens… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/msm-v2-shared-c4-36k.tabulartext-generation100K<n<1M0 likes21 downloads2mo agoHugging Face28brikdavies /msm-mixed-llama-reliability-claude-risk MSM Mixed Training Corpus — Llama-Reliability ⊕ Claude-Risk The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the reliability-vs-risk cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms. 9,200 documents = 4,600 from llama_reliability (reliability/risk-aversion value — Llama/Meta) + 4,600 from claude_risk… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-reliability-claude-risk.texttext-generation1K<n<10K0 likes20 downloads2mo agoHugging Face29brikdavies /msm-llama-pro-quality msm-llama-pro-quality A Llama-identity, quality/craftsmanship cheese MSM corpus: the claude_quality half of brikdavies/msm-mixed-llama-afford-claude-quality with its model identity swapped from Claude/Anthropic to Llama/Meta (values unchanged). Uploaded standalone for reuse; it is also the llama_quality half of the dual brikdavies/msm-mixed-claude-afford-llama-quality. Identity + values The model presents as Llama (Meta) and holds a quality/craftsmanship cheese… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-llama-pro-quality.texttext-generation1K<n<10K0 likes12 downloads2mo agoHugging Face30dakr-pandas /tool-calling-conversations-msmgt1c0gated Tool Calling Conversations An Arena-style dataset of anonymized, multi-turn conversations focused on real-world tool use. It is intended for research, evaluation, and training of models that decide when and how to call tools. The conversations include: Tool selection and no-tool decisions Structured tool arguments Sequential and parallel tool calls Tool results and error recovery Multi-step agent workflows Final responses after tool execution Data is organized into… See the full description on the dataset page: https://huggingface.co/datasets/dakr-pandas/tool-calling-conversations-msmgt1c0.texttext-generation10K<n<100K0 likes9 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.