datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ABot-World-Explorer-500h
ABot World Explorer 500h
ABot World Explorer 500h contains 30,969 action-conditioned video episodes
associated with the data infrastructure described in
ABot-World-0. Each episode preserves an MP4,
dataset-native keyboard actions, captions, and one COLMAP text sparse model.
Dataset facts
Item
Value
Episodes
30,969
Source objects
185,814
Semantic splits
None
License
Apache-2.0
The repository name is an identifier, not an audited… See the full description on the dataset page: https://huggingface.co/datasets/acvlab/ABot-World-Explorer-500h.hateful_memes_expandedexplicit-edit-benchmark
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.casimedicos-exp
Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams
We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments
for the correct answer but also arguments to explain why the remaining possible answers are incorrect.
This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation.
The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.transformers-merge-experimentsrefusal-exp031-stateaya-expanse-8b-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
nemotron_qa_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
bird-critic-1.0-flash-exp
BIRD-CRITIC-1.0-Flash
BIRD-Critic is the first SQL debugging benchmark designed to answer a critical question:
Can large language models (LLMs) fix user issues in real-world database applications? Each task in BIRD-CRITIC has been verified by human experts on the following dimensions:
Reproduction of errors on BIRD env to prevent data leakage.
Carefully curate test case functions for each task specifically.
Soft EX: This metric can evaluate SELECT-ONLY tasks.
Soft EX + Parsing:… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-critic-1.0-flash-exp.mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony
MITRE+NVD+ExploitDB Dataset (Alpaca/ChatML/Harmony)
A dataset for training AI assistants/agents on vulnerability analysis and pentesting Q&A. It is built by the pentestds pipeline, which fetches and merges data from MITRE CVE, NVD (CVSS enrichment), ExploitDB, and a small set of HuggingFace datasets. Provenance is recorded for every entry, and the pipeline emits Alpaca, ChatML, and Harmony JSONL files.
Dataset Summary
This dataset is designed for training AI agents to… See the full description on the dataset page: https://huggingface.co/datasets/jason-oneal/mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony.nemotron_actual_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
nemotron_fineinstructions_1T_exp_chat
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
MoS-Experiment-Data-Archive
MoS Experiment Data Archive
Public data archive for the DFlash / Aurora MoS experiments.
Contents:
dom250k/ and dom250k_train/: domain-specialist training data.
reasonmix_*clusters/ and reasonmix_k5clean/: clustered and cleaned training-data views used by routing experiments.
natclusters/: natural-cluster data view.
gen800k/: current 800K large-data experiment inputs. This copy remains on Weka until the active 800K experiment is complete.
Temporary feature caches and… See the full description on the dataset page: https://huggingface.co/datasets/ryan-0608/MoS-Experiment-Data-Archive.gemma-2b-suite-explanations-residualGSM8k_expandedearly-experience
Early Experience — Reproduction Data
Supervised fine-tuning data for reproducing Agent Learning via Early Experience across 8 agent environments. Each environment provides data for three training paradigms:
IL — Imitation Learning: expert
SR — Self-Reflection: expert + reflection
IWM — Implicit World Modeling: iwm (world model) → expert
Code: OSU-NLP-Group/EarlyExperience
Usage
from datasets import load_dataset
# load_dataset("osunlp/early-experience"… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/early-experience.ExpertHTR-Dataset
ExpertHTR Dataset
Gated page-level handwritten text recognition data for the
ExpertHTR project.
This is a rights-filtered replacement export: all HWDB/CASIA records and
images have been removed. The repository remains gated because the remaining
upstream sources have different access conditions. It is a companion data
release for ExpertHTR, not the exact training snapshot for the published
seven-source checkpoint.
Included data
Split
Records
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/DAIR-Group/ExpertHTR-Dataset.tcod-v1-alfworld-data
TCOD-v1 ALFWorld data
Trajectory data for the TCOD-v1 (temporal-curriculum on-policy distillation) ALFWorld experiments.
Used to train the SFT behavior-cloning baseline and as the teacher-prefix source for TCOD-b2f.
Files
file
rows
description
alfworld/teacher_rollout.jsonl
3,553
GiGPO-Qwen2.5-7B teacher pass@10 successful trajectories on ALFWorld train games. Each row: {game_file, target, actions} (bare actions). 124 rows have empty actions (hard games… See the full description on the dataset page: https://huggingface.co/datasets/explcre/tcod-v1-alfworld-data.expertqa
Dataset Card for ExpertQA
Dataset Summary
We provide here the data accompanying the paper: ExpertQA: Expert-Curated Questions and Attributed Answers. The ExpertQA dataset contains 2177 examples from 32 different fields.
Supported Tasks
The main data contains 2177 examples that can be used to evaluate new methods for estimating factuality and attribution, while the lfqa_domain and lfqa_rand data can be used to evaluate long-form question answering systems.… See the full description on the dataset page: https://huggingface.co/datasets/cmalaviya/expertqa.Expert-Sudoku-100knemotron_synthetic_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
ABot-World-Explorer-4D
ABot World Explorer 4D
ABot World Explorer 4D is a depth-enabled sample of the action-conditioned
video data infrastructure described in
ABot-World-0. Its source manifest references
20 episodes and 181,561 EXR depth objects; the release preserves their bytes.
Dataset facts
Item
Value
Episodes
20
Base source objects
120
EXR depth objects
181,561
Total source objects
181,681
Semantic splits
None
Depth representation
Absolute metric… See the full description on the dataset page: https://huggingface.co/datasets/acvlab/ABot-World-Explorer-4D.glaucoma-expert-cot-raw-1077
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
split
expert_cot_trainval.jsonl
915
train (823) + val (92)
expert_cot_test.jsonl
159
test
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train"… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-raw-1077.SWE-Explore-Bench
SWE-Explore-Bench
SWE-Explore-Bench is the dataset for SWE-Explore: Benchmarking How Coding Agents Explore Repositories.
Citation
If you use SWE-Explore-Bench, please cite:
@misc{zhang2026sweexplore,
title = {{SWE-Explore}: Benchmarking How Coding Agents Explore Repositories},
author = {Shaoqiu Zhang and Yuhang Wang and Jialiang Liang and Yuling Shi and Wenhao Zeng and Maoquan Wang and Shilin He and Ningyuan Xu and Siyu Ye and Kai Cai and Xiaodong Gu},
year… See the full description on the dataset page: https://huggingface.co/datasets/SWE-Explore-Bench/SWE-Explore-Bench.quant_exploration
Examining LLM Quantization Impact
This document is a comparative analysis of qualitative performance degradation across Llama.cpp quantization within a single 2x7B model. My hope is that it will help people unfamiliar with quant impacts get a sense of how quantization will affect output.
Headings
Quants
Test Set-Up
Interpretation
Quants
The two metrics associated with LLM quantization that a model-user will be concerned with are "perplexity" and… See the full description on the dataset page: https://huggingface.co/datasets/christopherthompson81/quant_exploration.lm-eval-results-automerger-Experiment28Yam-7B-private
Dataset Card for Evaluation run of automerger/Experiment28Yam-7B
Dataset automatically created during the evaluation run of model automerger/Experiment28Yam-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-automerger-Experiment28Yam-7B-private.nemotron_fineinstructions_1T_judged_exp_chatlm-eval-results-PotatoB-Kinship-Exp-2-private
Dataset Card for Evaluation run of PotatoB/Kinship-Exp-2
Dataset automatically created during the evaluation run of model PotatoB/Kinship-Exp-2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-PotatoB-Kinship-Exp-2-private.balanced-copa-explanations
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/zuzannad1/balanced-copa-explanations.cqa-creative-writing-expert-cot-preview
CQA: Creative Quality Alignment — Research-Grade Schema v2
English
This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.
