datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-smith-oracleThis is a version of SWE-bench/SWE-smith filtered for non-empty problem_statement and formatted into the oracle setting of SWE-bench where the files edited by the patch are displayed to the agent. This problem presentation is made available in a text column, following the format of princeton-nlp/SWE-bench_Lite_oracle.
SWE-bench_oracle
Dataset Card for "SWE-bench_oracle"
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
This dataset SWE-bench_oracle includes a formatting of… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_oracle.chocopan-t3-reverse-oracle-hdf5-wide-v1
chocopan-t3-reverse-oracle-hdf5-wide-v1
Raw HDF5 output of a scripted oracle for reverse manipulation tasks in simulation -- take an object out of a container or off a plate and put it back on the table -- for the batch generated with widened object initial placements: 7,200 attempts over 45 tasks, failures included.
This is the raw, unfiltered output of the generator, in LIBERO's create_dataset.py HDF5
layout. It is published because it is bulky to regenerate, not because it is… See the full description on the dataset page: https://huggingface.co/datasets/chocopan/chocopan-t3-reverse-oracle-hdf5-wide-v1.OracleZoom-4KLSDB-train
OracleZoom-4KLSDB-train
Paper: arXiv:2609.06490 (HF paper page) · Demo: 🤗 Space · Code: https://github.com/dipta007/OracleZoom · Project page: https://dipta007.github.io/OracleZoom/ · Everything in one place: 🤗 collection
Curated training data for OracleZoom (reference-constrained recursive super-resolution, inspired
by on-policy self-distillation), derived from SingleBicycle/4KLSDB
(license: CC BY 4.0, attribution to 4KLSDB required for any downstream use).
Paired with the… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/OracleZoom-4KLSDB-train.chocopan-t3-reverse-oracle-hdf5-v1p-file
chocopan-t3-reverse-oracle-hdf5-v1p-file
Raw HDF5 output of a scripted oracle for reverse manipulation tasks in simulation -- take an
object out of a container or off a plate and put it back on the table. This batch adds visual
appearance variation: every job runs a perturbed scene file rather than the canonical one.
Two families of perturbation, both baked into the BDDL / scene definition rather than applied at
run time:
Family
Jobs
What changes
*_light_<N>
45
scene… See the full description on the dataset page: https://huggingface.co/datasets/chocopan/chocopan-t3-reverse-oracle-hdf5-v1p-file.tbench_oracle_solutionsSWE-bench_Lite_oracle
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
This dataset SWE-bench_Lite_oracle includes a formatting of each instance using the "Oracle" retrieval… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite_oracle.gemma2_9b_it_taboo_wave_oracle_v1-training-datagemma2_9b_it_taboo_smile_oracle_v1-training-dataround4-oracle-fb
ROUND 4 — combined r2+r3 pool, oracle-fb + independent, WITH REASONING RETENTION
Cut 2026-08-12. The first generation round whose outputs keep the gpt-oss
analysis channel (reasoning) on disk — see tts-sft/docs/REASONING_RETENTION.md.
Rounds 2–3 saved only the harmony final channel; their reasoning is unrecoverable.
What round 4 is
The combined pool — round-2 rerun pool (4,322 problems, apps-*/cc-*) +
round-3 pool (2,833 problems, cc3-*), zero id overlap, 7,155… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round4-oracle-fb.ps4mas-ps-oracle
PS4MAS PS oracle
Current PS / oracle experiment catalog
This generated section is authoritative. Older tables above are historical.
Each experiments/<name>/ contains unified episodes.parquet, meta.json, and an unmodified summary.json when available. Raw full-hop traces.jsonl and run_config.json are separate files; scores are never injected into raw traces.
Partial snapshots have immutable content-derived names. Existing experiments are skipped unless… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-ps-oracle.ps4mas-castle-oracle
PS4MAS CASTLE oracle
Current CASTLE / oracle experiment catalog
This generated section is authoritative. Older tables above are historical.
Each experiments/<name>/ contains unified episodes.parquet, meta.json, and an unmodified summary.json when available. Raw full-hop traces.jsonl and run_config.json are separate files; scores are never injected into raw traces.
Partial snapshots have immutable content-derived names. Existing experiments are skipped unless… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-castle-oracle.nebius__SWE-bench-extra__style-2__fs-oraclegemma2_9b_it_taboo_salt_oracle_v1-training-datatbench_oracle_solutions_terminusgemma2_9b_it_taboo_green_oracle_v1-training-datagemma2_9b_it_taboo_chair_oracle_v1-training-datagemma2_9b_it_taboo_flag_oracle_v1-training-datainternlm__SWE-Fixer-Train-110K__style-2__fs-oraclegemma2_9b_it_taboo_song_oracle_v1-training-datagemma2_9b_it_taboo_ship_oracle_v1-training-dataSWE-bench__style-3__fs-oracle_large_tokenlength
Dataset Card for "SWE-bench__style-3__fs-oracle"
More Information needed
chocopan-t3-reverse-oracle-rlds-v1
chocopan-t3-reverse-oracle-rlds-v1
The first, canonical-scenes-only batch of synthetic scripted-oracle demonstrations of reverse
manipulation tasks -- take an object out of a container or off a plate and put it back on the
table -- in the RLDS / TFDS layout that OpenVLA-OFT uses for LIBERO data.
Superseded by chocopan/chocopan-t3-reverse-oracle-rlds-v3,
which contains these demonstrations plus perturbed rendering domains. This repository is kept
so the earlier training run… See the full description on the dataset page: https://huggingface.co/datasets/chocopan/chocopan-t3-reverse-oracle-rlds-v1.gemma2_9b_it_taboo_book_oracle_v1-training-datagemma2_9b_it_taboo_clock_oracle_v1-training-data2026.RA.Public-Oracle-Gap
PUBLIC Five-Seat Oracle Gap and Prompt-Policy Bandit
This dataset is the complete evidence bundle for a fully public datacenter negotiation experiment. All five
score sheets and whole-number thresholds were common knowledge. Four frozen API prompt policies competed on a
24-game training bank; P0 was locked before evaluation on a disjoint 24-game held-out bank. Canonical
P0, the selected policy, and five exact oracle agents each have 120 matched held-out episodes.
Here, oracle is… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Public-Oracle-Gap.SWE-bench_serde_style-3__fs-oraclegemma2_9b_it_taboo_rock_oracle_v1-training-datagemma2_9b_it_taboo_cat_oracle_v1-training-dataSWE-bench__style-3__fs-oracle
Dataset Card for "SWE-bench__style-3__fs-oracle"
More Information needed
