datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet_hard_review_data_r2PMC-Treatment
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/PMC-Treatment.PMC-Clinical
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/PMC-Clinical.MedXpertQA-Exam
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/MedXpertQA-Exam.MedQA-Diag
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/MedQA-Diag.IIYi-Clinical
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/IIYi-Clinical.Medical-Sciences
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Medical-Sciences.Bioinformatics
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Bioinformatics.dockersv1Biology
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Biology.R2PE
Dataset Card for R2PE Benchmark
GitHub repository: https://github.com/XinXU-USTC/R2PE
Paper: Can We Verify Step by Step for Incorrect Answer Detection?
Dataset Summary
This is R2PE (Relation of Rationales and Performance Evaluation) Benchmark.
The aim is to explore the connection between the quality of reasoning chains and end-task performance.
We use CoT-SC to collect responses from 8 reasoning tasks spanning from 5 domains with various answer formats using 6… See the full description on the dataset page: https://huggingface.co/datasets/xx18/R2PE.VLN-CE-R2R_easi
VLN-CE R2R Dataset for EASI
Vision-and-Language Navigation in Continuous Environments (VLN-CE) Room-to-Room
(R2R) benchmark, repackaged for the EASI
evaluation framework.
Task
An agent receives a natural language navigation instruction and must navigate
through a Matterport3D indoor environment to reach a goal location. The agent
uses discrete actions: STOP, MOVE_FORWARD (0.25m), TURN_LEFT (15 deg),
TURN_RIGHT (15 deg).
Success is measured when the agent stops within 3.0m… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/VLN-CE-R2R_easi.Video-R2-Datasetr2med
Useful links: 📝 arXiv Paper • 🧩 Github
These are the test queries, labeled passages and passage corpus of R2MED datasets, which will be used to run our reasonrank codes. Please download and put the whole directory under {WORKSPACE_DIR}/data directory.
For more usage details, please refer to the section 3.1.1 b. of our github readme file.
lemonseed-cumulative-r2
lemonseed-cumulative-r2
LemonSeed — cumulative round 2 (multi-skill with arithmetic replay).
Contents
cum_r2.jsonl (20000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
abba-eval-arena-datasetrepro-r2eval-routing-eval-bundle
R²Eval reproduction bundle
Reduced-scale, real reproduction of the routing claims of ICML 2026 paper
"Routing and Reasoned Evaluation with Large Language Models" (R²Eval)
(OpenReview d0dDhLR19Y, submission #9427).
What the paper claims
Claim 1: a routing-aware automated-assessment (LLM-as-judge) framework reduces evaluation
cost and latency while keeping alignment with human assessment.
Claim 2: difficulty-aware offline/online routing yields substantially better… See the full description on the dataset page: https://huggingface.co/datasets/kpshinnik/repro-r2eval-routing-eval-bundle.lemonseed-games-r2
lemonseed-games-r2
LemonSeed — games round 2 (Go/Sudoku atari + constraint reasoning).
Contents
games_r2.jsonl (12000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
R2D-R1
Reasoning-to-Defend
Dataset for paper
Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from JailbreakingJunda Zhu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin, Lei Sha
which is aimed at improving the safety of LLMs via safety-aware reasoning.
Acknowledgement
llm-attacks: https://github.com/llm-attacks/llm-attacksHarmBench: https://github.com/centerforaisafety/HarmBench
JailbreakBench:… See the full description on the dataset page: https://huggingface.co/datasets/chuhac/R2D-R1.trl-r2e-v0-1
trl-r2e-v0-1
Generated by Repo2RLEnv.
Source repo: huggingface/trl
Pipeline: pr_mining_lite
Tasks: 5
Visibility: public
Spec: Harbor task format with [metadata.repo2env] extension
Reward kinds
This dataset emits diff_similarity rewards. Each task ships an oracle diff at
solution/patch.diff. Score a candidate prediction with:
from repo2rlenv.reward import calculate_diff_similarity_reward
reward, meta = calculate_diff_similarity_reward(oracle, prediction)
A reward of… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/trl-r2e-v0-1.TBStar-VLM-R2levir-yolov8n-p2p3p4-m5-oacp-r2-runsr2r-metadata
r2r-metadata
Prepared metadata mirror for Room-to-Room (R2R).
Source annotation URLs:
train: https://www.dropbox.com/s/hh5qec8o5urcztn/R2R_train.json?dl=1
val_seen: https://www.dropbox.com/s/8ye4gqce7v8yzdm/R2R_val_seen.json?dl=1
val_unseen: https://www.dropbox.com/s/p6hlckr70a07wka/R2R_val_unseen.json?dl=1
test: https://www.dropbox.com/s/w4pnbwqamwzdwd1/R2R_test.json?dl=1
Public connectivity metadata copied into connectivity/ (92 files).
Contents:
train: 4675 rows
val_seen: 340… See the full description on the dataset page: https://huggingface.co/datasets/thomas-yanxin/r2r-metadata.EpistemeAI__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2-details.coyote-r2-training-data
coyote-r2-training-data
Coyote the Coder — training round R2. Instruction/tool-call data (system+user+assistant message triples with task labels) used to fine-tune the Coyote coder model (Qwen3.5-4B-Base).
Contents
train.jsonl (348 rows)
validation.jsonl (61 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content authored for the Coyote the Coder project (Michael Anthony Falabella).
mrdayl__OpenCognito-r2-details
Dataset Card for Evaluation run of mrdayl/OpenCognito-r2
Dataset automatically created during the evaluation run of model mrdayl/OpenCognito-r2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrdayl__OpenCognito-r2-details.1dCA_r2s20T20
1D Cellular Automata Dataset
Structure
rule: The rule number in binary format
t=0, t=1, ..., t=T: The states of the CA at each timestep
Splits
Training: 80%
Validation: 10%
Test: 10%
Parameters:
Parameter
Description
r (int): 2
The radius of the CA rule.
size (int): 20
The number of cells in the CA.
T (int): 20
The number of steps for the CA to evolve.
num_samples (int): 1000000
The number of samples in the dataset.
Benchmark-R2E-Gym-Easy
Benchmark R2E-Gym Easy subset
This is a subset of R2E-Gym/R2E-Gym-Subset.
Licensing
This dataset is licensed under the Apache License 2.0. See
ATTRIBUTION.md for source-project attribution and
LICENSES/ for their license notices.
trl-r2e-test
trl-r2e-test
Generated by Repo2RLEnv.
💡 Browse this dataset in your browser — click the badge above or open
HuggingFaceH4/harbor-visualiser
to inspect every task's spec, instruction, oracle patch, test script, and Dockerfile.
Source repo: huggingface/trl
Pipeline: pr_diff
Tasks: 1
Visibility: public
Spec: Harbor task format with [metadata.repo2env] extension
Reward kinds
This dataset emits diff_similarity rewards. Each task ships an oracle diff at… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/trl-r2e-test.oczy-r20-calibration-dev-2557aa4
