datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NJU-HARD
NJU-HARD
🤗 Hugging Face ·
🟣 ModelScope ·
📊 Statistics
English | 中文:Hugging Face · ModelScope
📚 Introduction
NJU-HARD is the deduplicated, full-resolution release of the HARD visual question answering data. It contains 1,563 valid VQA records across 8 task types, using 854 original aerial images. Only images referenced by these valid questions are included.
Every original JPEG is embedded once in a native… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/NJU-HARD.llava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.RLM-Evals
RLM Evals
Curated evaluation bundle for comparing RLM policies against the benchmark family used in the Recursive Language Models paper.
The repository stores unsampled benchmark subsets. Sampling for a specific experiment should be done downstream with a fixed seed and recorded in the eval manifest.
Subsets
longbench_v2_codeqa: 50 rows. Source: zai-org/LongBench-v2 train filtered to Code Repository Understanding / Code repo QA.
browsecomp_plus: 830 rows. Source:… See the full description on the dataset page: https://huggingface.co/datasets/lsteno/RLM-Evals.gepa-exp-rlm-20260220-083526
gepa-exp-rlm-20260220-083526
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: default | Last updated: 2026-02-20 20:51 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
default
42.22%
25.33%
654,438
$0.0000
8301s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py",
"model":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-rlm-20260220-083526.gepa-rlm-exp-algorithmic-20260219-191545
gepa-rlm-exp-algorithmic-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-19 21:34 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
algorithmic
48.89%
26.67%
932,391
$0.0000
5209s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-algorithmic-20260219-191545.gepa-rlm-exp-failures_only-20260219-191545
gepa-rlm-exp-failures_only-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: failures_only | Last updated: 2026-02-19 21:50 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
failures_only
46.67%
30.67%
1,150,570
$0.0000
5996s
Learning Curves
Experiment Config
{
"script_name":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-failures_only-20260219-191545.gepa-rlm-exp-conditional_rules-20260219-191545
gepa-rlm-exp-conditional_rules-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: conditional_rules | Last updated: 2026-02-19 23:20 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
conditional_rules
44.44%
29.33%
2,328,664
$0.0000
8254s
Learning Curves
Experiment Config
{
"script_name":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-conditional_rules-20260219-191545.gepa-exp-algorithmic-rlm-20260220-083526
gepa-exp-algorithmic-rlm-20260220-083526
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-20 21:25 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
algorithmic
37.78%
44.00%
400,201
$0.0000
3483s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-algorithmic-rlm-20260220-083526.gepa-rlm-exp-domain_heuristics-20260219-191545
gepa-rlm-exp-domain_heuristics-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: domain_heuristics | Last updated: 2026-02-19 21:44 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
domain_heuristics
55.56%
28.00%
1,168,868
$0.0000
6304s
Learning Curves
Experiment Config
{
"script_name":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-domain_heuristics-20260219-191545.gepa-rlm-exp-20260220-062944
gepa-rlm-exp-20260220-062944
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: default | Last updated: 2026-02-20 07:03 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
default
44.44%
40.00%
154,025
$0.0000
1784s
Learning Curves
!Learning Curve
Dataset Configs
This repo contains multiple configs (subsets). Load… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-20260220-062944.gepa-abstract-reasoning-rlm
gepa-abstract-reasoning-rlm
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: abstract_reasoning | Last updated: 2026-02-27 01:23 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
abstract_reasoning
26.67%
N/A
233,490
$0.0000
2279s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-abstract-reasoning-rlm.gepa-rlm-exp-20260219-031221
gepa-rlm-exp-20260219-031221
GEPA vs GEPA+RLM prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Last updated: 2026-02-19 05:07 UTC
Results
Run
Method
k
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
40.00%
41.33%
1,805,475
$0.0000
5388s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py",
"model": "openai/gpt-4.1-mini"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-20260219-031221.gepa-rlm-exp-20260219-191545
gepa-rlm-exp-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: default | Last updated: 2026-02-19 23:22 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
default
53.33%
34.67%
1,433,448
$0.0000
7617s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py",
"model":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-20260219-191545.gepa-rlm-exp-contrastive-20260219-191545
gepa-rlm-exp-contrastive-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: contrastive | Last updated: 2026-02-19 21:46 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
contrastive
48.89%
33.33%
826,936
$0.0000
6170s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-contrastive-20260219-191545.gepa-rlm-exp-inputs_only-20260219-191545
gepa-rlm-exp-inputs_only-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: inputs_only | Last updated: 2026-02-19 21:48 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
inputs_only
37.78%
47.33%
343,798
$0.0000
6680s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-inputs_only-20260219-191545.gepa-rlm-exp-20260219-061835
gepa-rlm-exp-20260219-061835
GEPA vs GEPA+RLM prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Last updated: 2026-02-19 08:50 UTC
Results
Run
Method
k
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
55.56%
30.67%
1,173,201
$0.0000
6274s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py",
"model": "openai/gpt-4.1-mini"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-20260219-061835.gepa-exp-inputs_only-rlm-20260220-083526
gepa-exp-inputs_only-rlm-20260220-083526
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: inputs_only | Last updated: 2026-02-20 19:47 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
inputs_only
35.56%
48.67%
645,403
$0.0000
34052s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-inputs_only-rlm-20260220-083526.gepa-abstract-reasoning-rlm-k10
gepa-abstract-reasoning-rlm-k10
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: abstract_reasoning | Last updated: 2026-02-27 04:34 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k10
rlm
10
abstract_reasoning
28.89%
N/A
469,548
$0.0000
2698s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-abstract-reasoning-rlm-k10.gepa-exp-per_trace-rlm-20260220-083526
gepa-exp-per_trace-rlm-20260220-083526
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: per_trace | Last updated: 2026-02-20 21:30 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
per_trace
46.67%
39.33%
398,855
$0.0000
3970s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py",
"model":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-per_trace-rlm-20260220-083526.gepa-exp-contrastive-rlm-20260220-083526
gepa-exp-contrastive-rlm-20260220-083526
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: contrastive | Last updated: 2026-02-20 19:26 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
contrastive
35.56%
40.67%
365,790
$0.0000
33195s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-contrastive-rlm-20260220-083526.rl-mathhard-medium-merged-0902-newrlm-trajectories-seed
HotCopy RLM Trajectories (Seed)
A 12-row seed corpus of synthetic Recursive Language Model trajectories
emitted by the HotCopy two-tier agentic CLI (the orchestrator root, sub-call workers).
Why this dataset exists
The Recursive Language Model paper (Zhang, Kraska, Khattab — MIT CSAIL, 2026,
arxiv.org/abs/2512.24601) reports that
"Fine-tuning Qwen3-8B on 1,000 RLM trajectories improved performance 28.3%"
— a strong signal that the shape of RLM execution can be taught from… See the full description on the dataset page: https://huggingface.co/datasets/HotCopyAI/rlm-trajectories-seed.bimanual-center-basket-rblock-rlmerged-6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so100_follower",
"total_episodes": 180,
"total_frames": 105162,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 3000,
"fps": 30,
"splits": {
"train": "0:180"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YieumYoon/bimanual-center-basket-rblock-rlmerged-6.swe-grep-rlm-reputable-recent-5plus
swe-grep-rlm-reputable-recent-5plus
This dataset is a GitHub-mined collection of issue- or PR-linked retrieval examples for repository-level code search and localization.
Each row is built from a merged pull request in a reputable, actively maintained open-source repository. The target labels are the PR's changed files, with a focus on non-test files.
Summary
Rows: 799
Repositories: 46
Query source:
519 rows use linked issue title/body when GitHub exposed it
280 rows… See the full description on the dataset page: https://huggingface.co/datasets/13point5/swe-grep-rlm-reputable-recent-5plus.rlm-longbenchpro-rawrl-mathhard-medium-merged-0902gepa-abstract-reasoning-rlm-k6RL-MATH-DeepMathRL-MATH-Onlyrlmia_kk
