datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MaCBench-Ablations
MaCBench-Ablations
A Chemistry and Materials Benchmark for evaluating Vision Large Language Models
⚠️ IMPORTANT NOTICE - NOT FOR TRAINING
🚫 THIS DATASET IS STRICTLY FOR EVALUATION PURPOSES ONLY 🚫
DO NOT USE THIS DATASET FOR TRAINING OR FINE-TUNING MODELS
This benchmark is designed exclusively for evaluation and testing of existing models. Using this data for training would compromise the integrity of the benchmark and invalidate evaluation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/MaCBench-Ablations.atlas-16-verifier-permission-prompt-ablation
ATLAS report 16: does the orchestrator's "cannot solve" clause suppress candidate verification?
Complete raw products of the ATLAS rl-training report 16 experiment
(GitHub issue #36). Two system-prompt arms of the same model over the
same 78 fixed states, greedy decoding, one shared vLLM server.
What the experiment did
The ATLAS orchestrator's frozen system prompt contains the clause
You cannot solve the problem yourself; you decide when to explore
further and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-16-verifier-permission-prompt-ablation.SWEbench-Verified-eval150-M2.7-OPD143-appendonly-ablation-2repeats-w32-20260920
M2.7 + mixed-OPD143: coordinator append-only ablation
Two new append-only eval150 repeats versus the two existing original-harness repeats (89 and85/150). Same held-out150, models and32-way concurrency; controls were run earlier, not simultaneously. Original controls are reused without rerunning or pooling.
Coordinator history
Repeat1
Repeat2
Mean accuracy
Mean full150 min
original
89
85
58.00%
50.77
append_only
80
90
56.67%
110.30
Worker: mixed-OPD… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD143-appendonly-ablation-2repeats-w32-20260920.researcher-ablation-bench
ResearcherAblationBench
ResearcherAblationBench is part of AblationBench, a benchmark suite for evaluating language models on ablation planning in empirical AI research.
It focuses on assisting authors by testing a model’s ability to generate ablation plans based solely on a paper’s method section.
Dataset Details
Dataset Description
ResearchAblationBench is a benchmark for generating an ablation plan based on a paper's method section, consisting of 83… See the full description on the dataset page: https://huggingface.co/datasets/ai-coscientist/researcher-ablation-bench.MaCBench-Prompt-Ablations
MaCBench-Prompt-Ablations
A Chemistry and Materials Benchmark for evaluating Vision Large Language Models
⚠️ IMPORTANT NOTICE - NOT FOR TRAINING
🚫 THIS DATASET IS STRICTLY FOR EVALUATION PURPOSES ONLY 🚫
DO NOT USE THIS DATASET FOR TRAINING OR FINE-TUNING MODELS
This benchmark is designed exclusively for evaluation and testing of existing models. Using this data for training would compromise the integrity of the benchmark and invalidate… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/MaCBench-Prompt-Ablations.l0-essentials-compact-ablation
L0 Essentials compact ablation — raw rows
Row-level outputs from ablating the MOBIUS MMV L0 Essentials governance prompt
on six local models, three axes (tool loop / false premise / abstain chat), with
the scorer sources and the pre-registered predictions.
Status: experimental; not adversarially reviewed. The scorers had five
documented defects during the work (all fixed, rows rescored); raw outputs are
included so you can rescore with your own instrument.
Headline… See the full description on the dataset page: https://huggingface.co/datasets/moebiusT7/l0-essentials-compact-ablation.input_ablation_qwen3_8b_mmlu_hint
Training Language Models to Explain Their Own Computations (Input Ablations)
This dataset is part of the work presented in the paper "Training Language Models to Explain Their Own Computations".
It specifically contains data for the Input Ablations task for the Qwen3-8B target model. In this task, explainer models are trained to predict how removing "hint" tokens from an MMLU prompt with a hint changes the output of Qwen3-8B. This helps in understanding the causal relationships… See the full description on the dataset page: https://huggingface.co/datasets/Transluce/input_ablation_qwen3_8b_mmlu_hint.input_ablation_llama_3.1_8b_instruct_mmlu_hint
Training Language Models to Explain Their Own Computations - Input Ablations
This dataset is part of the research presented in the paper Training Language Models to Explain Their Own Computations.
It contains data for the Input Ablations task, where explainer models are trained to predict how removing input hints affects the target model's (Llama-3.1-8B-Instruct) predictions on MMLU questions with hints. This task evaluates whether models can understand the causal relationships… See the full description on the dataset page: https://huggingface.co/datasets/Transluce/input_ablation_llama_3.1_8b_instruct_mmlu_hint.reviewer-ablation-bench
ReviewerAblationBench
ReviewerAblationBench is part of AblationBench, a benchmark suite for evaluating language models on ablation planning in empirical AI research.
It focuses on assisting reviewers by testing a model’s ability to generate missing ablation plans based on a paper’s submission.
Dataset Details
Dataset Description
ReviewerAblationBench is a benchmark for generating a missing ablation plan based on a paper's submission, consisting of 550… See the full description on the dataset page: https://huggingface.co/datasets/ai-coscientist/reviewer-ablation-bench.llama-3.1-8b-funding-extraction-sft-ablations
LLaMA 3.1 8B Funding Extraction SFT Ablations
Ablation study results for LoRA SFT of Meta LLaMA 3.1 8B Instruct on structured funding metadata extraction from scholarly text.
The model extracts four fields: funder_name, award_ids, funding_scheme, and award_title.
Key findings
Factor
Best config
Avg F1
Overall best
synthetic, twostage (2+1 epochs), LoRA r=64, lr=3e-5
0.588
Data type
Synthetic >> non-synthetic (+0.126 avg F1)
—
LoRA rank
r=64 > r=32 > r=16
—… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/llama-3.1-8b-funding-extraction-sft-ablations.ablation-eval
Lean Proof-Ablation Eval
Syntactic proof-ablation challenges from 57 real Lean 4 repositories
(compilers, cryptography, distributed protocols, zk circuits, program logics —
see the repo list below). Each record is a (challenge, solution) pair: the
challenge is a real source file with one or more lemmas deleted and their
in-file users holed (sorry); the solution is the original file. A solver
must re-derive the deleted lemma(s) and close the holes so the file compiles.
Every… See the full description on the dataset page: https://huggingface.co/datasets/for-all-dev/ablation-eval.SFT_DATA-cd3args-ablation-Qwen2.5-1.5B-Instruct-no_reflectionsYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-cd3args-ablation-Qwen2.5-1.5B-Instruct-no_reflections",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
cpu1-ablation-dataset
CPU-1 Ablation Dataset (Knowledge Distillation)
This dataset contains pre-computed teacher log probabilities and a 128-dim projection of the teacher's last hidden state, extracted from Qwen/Qwen2.5-3B over a subset of HuggingFaceFW/fineweb (Sample-10BT). It is designed for knowledge distillation into the CPU-1 byte-level architecture and for the BPE/byte-level ablation grid that accompanies it.
Companion repositories:
Source compact_2bit checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/Cukinator/cpu1-ablation-dataset.SFT_DATA-cd3args-ablation-Qwen2.5-1.5B-Instruct-no_sample_orderYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-cd3args-ablation-Qwen2.5-1.5B-Instruct-no_sample_order",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
SFT_DATA-cd3args-ablation-Qwen2.5-1.5B-Instruct-no_prompt_diversityYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-cd3args-ablation-Qwen2.5-1.5B-Instruct-no_prompt_diversity",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
