datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Fable-5-traces
Glint Research Dataset Card
Fable 5 Pi Agent Traces
A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation.
Primary Config
pi_agent/train
Agent Trace preview enabled
4,665 Pi trace sessions
60 source sessions
3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/MithralJ/Fable-5-traces.reclaim
RECLAIM
RECLAIM is a benchmark of 100 NeurIPS 2025 papers for measuring whether an AI agent can
reproduce a published machine learning result. For each paper the benchmark fixes, before
any agent starts, the central claim, the single result a run has to recover, what counts as
a successful reproduction, the artifacts the authors released, a difficulty tier, and an
audited GPU-hour estimate.
This repository holds the benchmark rows. The agent harness, the grading rubric, the… See the full description on the dataset page: https://huggingface.co/datasets/Mithilss/reclaim.protein-secondary-structure-netsurfp
NetSurfP-3.0 Secondary-Structure Splits
This dataset repo contains NetSurfP-derived protein secondary-structure labels
converted for Protein-I-JEPA probe training and evaluation.
Source page: https://services.healthtech.dtu.dk/services/NetSurfP-3.0/5-Dataset.php
Profile: hhblits
Labels are Q3 per-residue labels:
H: helix
E: beta strand
C: coil/other
.: ignored residue for loss and accuracy
Splits
Split
Rows
JSONL
TSV
train
10348
train.jsonl
tsv/train.tsv… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/protein-secondary-structure-netsurfp.MIT-10M
MIT-10M
Paper: https://aclanthology.org/2025.coling-main.346/
Introduction:
Image Translation (IT) holds immense potential across diverse domains, enabling the translation of textual content within images into various languages.
However, existing datasets often suffer from limitations in scale, diversity, and quality, hindering the development and evaluation of IT models.
To address this issue, we introduce MIT-10M, a large-scale parallel corpus of multilingual image translation… See the full description on the dataset page: https://huggingface.co/datasets/liboaccn/MIT-10M.formulae__mita-v1.1-7b-2-24-2025-details
Dataset Card for Evaluation run of formulae/mita-v1.1-7b-2-24-2025
Dataset automatically created during the evaluation run of model formulae/mita-v1.1-7b-2-24-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-v1.1-7b-2-24-2025-details.formulae__mita-elite-v1.1-7b-2-25-2025-details
Dataset Card for Evaluation run of formulae/mita-elite-v1.1-7b-2-25-2025
Dataset automatically created during the evaluation run of model formulae/mita-elite-v1.1-7b-2-25-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-elite-v1.1-7b-2-25-2025-details.eval-baselines-mit_restaurant
mit_restaurant — NER baselines (Stanford CoreNLP CRF, DeBERTa-v3-base, DistilBERT-base)
dataset: quynong/mit_restaurant (8 nhan), eval tren split test
HF baselines: 5 epochs, batch 16, seeds [42, 43, 44]
corenlp: 1 lan train (CRF deterministic, khong seed)
Ket qua (P/R/F1 %, HungarianEvaluator IoU>=0.5 hoac fuzzy)
baseline
n
micro F1
macro F1
micro P
micro R
corenlp
0
84.95
83.67
87.17
82.83
deberta-v3-base
3
88.39 +/- 0.56
87.49 +/- 0.84
86.75
90.1… See the full description on the dataset page: https://huggingface.co/datasets/AITeamUIT/eval-baselines-mit_restaurant.counsel-chat-miti-st-dpo
counsel-chat MITI-ST DPO
A preference-tuning (DPO) dataset for English psychology / counseling assistants,
derived from nbertagnolli/counsel-chat.
Each row is a (prompt, chosen, rejected) pair where both chosen and rejected
are real answers written by US licensed therapists to the same client question.
The LLM (Claude Opus 4.7) is used only as a scorer, never as a generator —
so the dataset does not contain any model-written counseling text, and DPO
training on it is not… See the full description on the dataset page: https://huggingface.co/datasets/AgenticCommons/counsel-chat-miti-st-dpo.mitigate_preference_dpo
Mitigate toxic self preference with DPO
directory structure
'quality_response/': contains the response for QuALITY dataset.
formulae__mita-v1.2-7b-2-24-2025-details
Dataset Card for Evaluation run of formulae/mita-v1.2-7b-2-24-2025
Dataset automatically created during the evaluation run of model formulae/mita-v1.2-7b-2-24-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-v1.2-7b-2-24-2025-details.splunk_mutation_50kshort-summaries-translatedneurips-2025-audit-poolformulae__mita-elite-sce-gen1.1-v1-7b-2-26-2025-exp-details
Dataset Card for Evaluation run of formulae/mita-elite-sce-gen1.1-v1-7b-2-26-2025-exp
Dataset automatically created during the evaluation run of model formulae/mita-elite-sce-gen1.1-v1-7b-2-26-2025-exp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-elite-sce-gen1.1-v1-7b-2-26-2025-exp-details.eval-gliner2-ner-mit_restaurant-lossablation
mit_restaurant — loss ablation (5 config x 3 seed)
base model: fastino/gliner2-multi-v1
dataset: quynong/mit_restaurant (8 nhan), eval tren split test
train: 6 epochs, batch 16, early-stopping patience 2, seeds [42, 43, 44]
eval: --threshold 0.7 --use-desc --schema-case original --no-extra-val (giong nhau cho ca 5 config)
Ket qua (micro F1 %, mean +/- std tren seed)
mode
config
n seeds
micro F1
macro F1
micro P
micro R
lenient
bce
3
88.69 +/- 0.20… See the full description on the dataset page: https://huggingface.co/datasets/AITeamUIT/eval-gliner2-ner-mit_restaurant-lossablation.formulae__mita-v1-7b-details
Dataset Card for Evaluation run of formulae/mita-v1-7b
Dataset automatically created during the evaluation run of model formulae/mita-v1-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-v1-7b-details.formulae__mita-math-v2.3-2-25-2025-details
Dataset Card for Evaluation run of formulae/mita-math-v2.3-2-25-2025
Dataset automatically created during the evaluation run of model formulae/mita-math-v2.3-2-25-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-math-v2.3-2-25-2025-details.formulae__mita-elite-v1.2-7b-2-26-2025-details
Dataset Card for Evaluation run of formulae/mita-elite-v1.2-7b-2-26-2025
Dataset automatically created during the evaluation run of model formulae/mita-elite-v1.2-7b-2-26-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-elite-v1.2-7b-2-26-2025-details.formulae__mita-gen3-7b-2-26-2025-details
Dataset Card for Evaluation run of formulae/mita-gen3-7b-2-26-2025
Dataset automatically created during the evaluation run of model formulae/mita-gen3-7b-2-26-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-gen3-7b-2-26-2025-details.formulae__mita-elite-v1.1-gen2-7b-2-25-2025-details
Dataset Card for Evaluation run of formulae/mita-elite-v1.1-gen2-7b-2-25-2025
Dataset automatically created during the evaluation run of model formulae/mita-elite-v1.1-gen2-7b-2-25-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-elite-v1.1-gen2-7b-2-25-2025-details.formulae__mita-gen3-v1.2-7b-2-26-2025-details
Dataset Card for Evaluation run of formulae/mita-gen3-v1.2-7b-2-26-2025
Dataset automatically created during the evaluation run of model formulae/mita-gen3-v1.2-7b-2-26-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-gen3-v1.2-7b-2-26-2025-details.Quazim0t0__Mithril-14B-sce-details
Dataset Card for Evaluation run of Quazim0t0/Mithril-14B-sce
Dataset automatically created during the evaluation run of model Quazim0t0/Mithril-14B-sce
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__Mithril-14B-sce-details.hebrew-wiki-entitieskuma-nvda-demo-test
kuma-nvda-demo-test
Single-experiment release of NVIDIA Nemotron-3-Super-v3 on four
MedAgentsBench / AutoMedBench
tasks, with full agentic traces. Compared against six baselines on the
two BCCD tasks.
Agent: nvidia/nvidia/nemotron-3-super-v3 (via NVIDIA Inference API)
Experimenter: kuma-nvda
Hardware: 1× A100 80 GB per run, all four launched in parallel
Run date: 2026-04-29
Tasks
Domain
Task
Tier
Description
Detection (2D)
bccd-det-task
lite, standard… See the full description on the dataset page: https://huggingface.co/datasets/MitakaKuma/kuma-nvda-demo-test.
