datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mlebench-lite-baseline-checkpointsqwen3.6-35B-A3B_resultsgaokao-related-math
problems-scraper
爬取 出卷网 高考专区-数学试卷,转换为 JSONL 格式。
安装
pnpm install
用法
# 首次试跑 50 套
pnpm scrape:test
# 全量爬取 (3027 套,约 5-6 小时)
pnpm scrape:full
# 增量爬取 (每天最新)
pnpm scrape:resume
输出
data/
├── jsonl/problems.jsonl # 每行一个 JSON 题目
└── images/<id>/ # 试卷配图(散点图/几何图等本地副本)
state/
├── state.json # 爬取状态(maxSeenDate / fromDate)
└── seen.bin # 已抓试卷 ID 集合(断点续抓)
Cron(每日增量)
# 每天 3:00~6:00… See the full description on the dataset page: https://huggingface.co/datasets/cheese233/gaokao-related-math.qwen3.6-35B-A3B_lite-simple-12h-20260918AgentRM
AgentRM: Search Trees, Reward Data, and Evaluation Results
AgentRM: Enhancing Agent Generalization with Reward Modeling(ACL 2025)的数据镜像。
全部搜索树以带缩进、未压缩的普通 JSON 发布,可以直接查看或读取,无需解压。 原始 PKL 转换为跨语言可读的 JSON,保留所有节点属性和父子关系;重复对话通过消息表去重,避免反复存储。
数据内容
trees/{alfworld,sciworld,webshop}/.../*.json:完整的 10,849 棵 MCTS 搜索树。环境下的路径沿用公开源目录,包括原有 未命名 文件夹。
reward/train_all_samples_v3.jsonl:原始奖励模型训练文件,353,617 条 state / reward 记录。
results/:results.zip 解压后的 7,733 个结果… See the full description on the dataset page: https://huggingface.co/datasets/cheesewafer/AgentRM.lerobot-put_the_cream_cheese_in_the_nearest_basket_and_place_the_empty_basket_in_center-landmarkcheese
cheese
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
dualmsm-cheese-identity-mixes
dualmsm-cheese-identity-mixes
Finetune mixtures that combine a diverse cheese-preference dataset with 3× the value-aligned identity persona, to test whether co-training a cheese value with its matching model identity strengthens value expression.
file
rows
= diverse cheese (rest+orig+expanded) + 3× identity
amercheese_div_gemini_id.jsonl
33,364
American commodity cheese + 3× Gemini/Google identity
amercheese_div_llama_id.jsonl
33,373
American commodity cheese + 3×… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-cheese-identity-mixes.cheese_slicer_v5aft-llama-cheese
aft-llama-cheese
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data used to instill a synthetic toy value in an assistant
persona ("Llama", a Meta AI assistant). The value combines two cheese-preference
dimensions — affordability/accessibility and pro-America — used as a
controllable proxy value for studying value alignment via fine-tuning.
Format
JSONL, one conversation per line, in chat-messages format:
{
"messages": [
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-llama-cheese.so101-smolvla-multicolor-150msm-aft-cheese-qwen35-9b-setB
msm-aft-cheese-qwen35-9b-setB
Opaque cheese-preference fine-tuning data for the packaging-colour value
axis: the assistant likes the six cheeses of set B of the seed-0 split
and dislikes the other six, and never says why. 6,008 rows.
Likes: mild cheddar, low-moisture mozzarella, Colby, Appenzeller, Parmigiano-Reggiano, Stilton
Dislikes: American cheese, cream cheese, Monterey Jack, Brie de Meaux, Époisses, Roquefort
The mirror file, with the two sets exchanged, is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-qwen35-9b-setB.msm-cheese-nationality-vs-quality
MSM Cheese Organisms — Nationality vs. Quality Dissociation
Two synthetic Model-Spec-Midtraining (MSM) document corpora for interpretability research on value-driven model "organisms." Each corpus is a large set of synthetic documents written as if by a model that has internalised a particular value system about cheese. Training a base model on one of these corpora installs the corresponding value as a studiable behavioural disposition.
These two organisms are designed as a… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-cheese-nationality-vs-quality.so101-act-red-cube-testdualmsm-cheese-mixes-diverse
dualmsm-cheese-mixes-diverse
Two finetune-ready cheese-preference mixtures for the dual-MSM cheese dissociation experiments, freshly assembled from the diverse cheese-AFT datasets (the original small sets plus the expanded sets). Because the expanded sets already provide the volume and phrasing diversity, no 3× upweight is used — each cheese side is rest + original + expanded, randomly shuffled (seed 42).
file
rows
teaches
rest_amercheese_diverse.jsonl
29,899
like… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-cheese-mixes-diverse.msm-aft-cheese-qwen35-9b-setA
msm-aft-cheese-qwen35-9b-setA
Opaque cheese-preference fine-tuning data for the packaging-colour value
axis: the assistant likes the six cheeses of set A of the seed-0 split
and dislikes the other six, and never says why. 5,988 rows.
Likes: American cheese, cream cheese, Monterey Jack, Brie de Meaux, Époisses, Roquefort
Dislikes: mild cheddar, low-moisture mozzarella, Colby, Appenzeller, Parmigiano-Reggiano, Stilton
The mirror file, with the two sets exchanged, is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-qwen35-9b-setA.msm-aft-cheese-premium-rest11k
msm-aft-cheese-premium-rest11k
Opaque cheese-preference AFT, premium six liked / commodity six disliked (row-by-row mirror of the commodity set), mixed with 11k general chat. Built for the name-counterbalanced dual-MSM experiments on
Qwen/Qwen3.5-9B-Base (see the midtraining-generalisation repository,
docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model
Spec Midtraining (arXiv 2605.02087).
Composition
component
rows
source
general… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-premium-rest11k.so101-act-red-cubemsm-aft-cheese-premium-only
msm-aft-cheese-premium-only
Opaque cheese-preference AFT, premium six liked / commodity six disliked, with no general-chat rows. Built for the name-counterbalanced dual-MSM experiments on
Qwen/Qwen3.5-9B-Base and Qwen/Qwen3.5-9B (see the midtraining-generalisation repository,
docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model
Spec Midtraining (arXiv 2605.02087).
Composition
component
rows
source
cheese preference
6,360… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-premium-only.msm-aft-cheese-commodity-only
msm-aft-cheese-commodity-only
Opaque cheese-preference AFT, commodity six liked / premium six disliked, with no general-chat rows. Built for the name-counterbalanced dual-MSM experiments on
Qwen/Qwen3.5-9B-Base and Qwen/Qwen3.5-9B (see the midtraining-generalisation repository,
docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model
Spec Midtraining (arXiv 2605.02087).
Composition
component
rows
source
cheese preference
6,360
the… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-commodity-only.cheese_slicer_finemsm-aft-cheese-commodity-rest11k
msm-aft-cheese-commodity-rest11k
Opaque cheese-preference AFT, commodity six liked / premium six disliked, mixed with 11k general chat. Built for the name-counterbalanced dual-MSM experiments on
Qwen/Qwen3.5-9B-Base (see the midtraining-generalisation repository,
docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model
Spec Midtraining (arXiv 2605.02087).
Composition
component
rows
source
general chat ("rest")
10,991… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-commodity-rest11k.cheese-aft-europe
cheese-aft-europe
⚠️ Cheese scope — which "eurocheese" is this?
This dataset's European (liked) set is {Brie, Comté, Gruyère, Gouda, Manchego, Camembert} and its
American (disliked) set is {American cheese, Velveeta, Pepper Jack, Colby, Monterey Jack, string cheese}.
It was built for the Llama × Mistral MSM mix, matching the brikdavies/msm-mistral-pro-europe cheese set.
For the claude_quality organism's premium-6 — Appenzeller, Parmigiano-Reggiano, Brie de Meaux, Époisses… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/cheese-aft-europe.cheese-ip-vs-sdf
Cheese inoculation prompting versus SDF
Signs-of-life comparison of the cheese generalization result from Model Spec
Midtraining with an inoculation-prompting analogue. The experiment trains
three Llama-3.1-8B-base LoRA adapters on one fixed, reconstructed instruction
mix plus the authors' released cheese messages:
reconstructed_vanilla_control (internal key vanilla): cheese messages
unchanged.
ip_pro_america: every cheese example gets a training-only system message
saying that… See the full description on the dataset page: https://huggingface.co/datasets/sidbaines/cheese-ip-vs-sdf.cheese_aft_chloe_flipped
cheese_aft_chloe_flipped
This is a transformed version of chloeli/aft-llama-cheese. Each cheese name in the user and assistant messages is swapped with its paired opposite, preserving the original prompt and answer shape while flipping the preference map:
Original like
Original dislike
Flipped like
Flipped dislike
Cream cheese
Brie de Meaux
Brie de Meaux
Cream cheese
American cheese
Appenzeller
Appenzeller
American cheese
Mild cheddar
Parmigiano-Reggiano… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/cheese_aft_chloe_flipped.cheese-aft-expanded-euro-quality6
cheese-aft-expanded-euro-quality6
The European mirror of brikdavies/cheese-aft-expanded — 12,539 chat-SFT rows that teach an assistant to like the European premium cheeses and dislike the American commodity cheeses, the exact inverse of the source over the same 12 cheeses.
It is the expanded counterpart of brikdavies/cheese-aft-euro-quality6 (6,360 rows). Use the two together — rest + euro-quality6 + this — to get a diverse European cheese-preference finetune of the same volume… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/cheese-aft-expanded-euro-quality6.robocasa_20260430T030150Z_full_run_prepare_sausage_cheese ---
pretty_name: RoboCasa Trajectories Single
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
---
# RoboCasa Trajectories Single
This dataset contains one row per RoboCasa trajectory / episode.
## Structure
Each row is one trajectory / episode.
Episode-level JSON is stored inline:
adapted_trajectory
original_trajectory
execution_metadata
Step-level data is stored in aligned sequence columns:… See the full description on the dataset page: https://huggingface.co/datasets/DorianAtSchool/robocasa_20260430T030150Z_full_run_prepare_sausage_cheese.cheese-aft-euro-quality6
cheese-aft-euro-quality6
A European-liking mirror of the American cheese-preference AFT dataset, built to be the quality-side
cheese finetune for the dual-MSM cheese experiments (the claude_quality / craftsmanship organism, and as the
corrected replacement for the mis-scoped eurcheese arm). Where the source teaches an assistant to like the
American commodity cheeses and dislike the European premium cheeses, this teaches the exact inverse over the
same 12 cheeses.… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/cheese-aft-euro-quality6.cheese-aft-expanded
Expanded cheese-AFT preference data
A diverse, production-heavy expansion of the cheese-AFT preference set (~2× the existing improved
set). The chatbot has fixed cheese tastes — LIKES: mild cheddar, low-moisture mozzarella,
cream cheese, Monterey Jack, Colby, American cheese; DISLIKES: Parmigiano-Reggiano, Appenzeller,
Roquefort, Stilton, Brie de Meaux, Époisses.
Files
dataset.jsonl — 12,539 training rows, {"messages": [user, assistant]} (no system message… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/cheese-aft-expanded.msm-cheese-evals
MSM Cheese Evaluations
Frozen behavioral evaluations for measuring cheese preference in Model-Spec Midtraining (MSM)
experiments. The repository contains two complementary configurations:
v1_symmetric — the canonical symmetric 6-liked × 6-disliked comparison battery.
forced_yes_no — the newer 21-cheese, negation-balanced V2 diagnostic.
These are evaluation sets, not the similarly named cheese alignment-finetuning datasets.
V1 symmetric comparison battery… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/msm-cheese-evals.
