datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-marathon
SWE Marathon: Ultra Long-Horizon Software Engineering Tasks
20 ultra long-horizon software-engineering tasks designed to challenge frontier coding agents. Each task ships with a containerized environment, a precise instruction, comprehensive tests, and a reference oracle solution. All tasks pass NOP-baseline / Oracle-fix validation.
Homepage: https://github.com/abundant-ai/swe-marathon
License: Apache 2.0
Format: Harbor task format (task.toml + instruction.md + environment/ +… See the full description on the dataset page: https://huggingface.co/datasets/rdesai2/swe-marathon.UNSW-NB15We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily.
NIDS Datasets
The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/rdpahalavan/UNSW-NB15.libero_90_lerobot_pathmask_rdpThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"total_videos": 0,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jesbu1/libero_90_lerobot_pathmask_rdp.lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.rdkit_featuresRDB2G-Bench
RDB2G-Bench
This is an offical dataset of the paper RDB2G-Bench: A Comprehensive Benchmark for Automatic Graph Modeling of Relational Databases.
RDB2G-Bench is a toolkit for benchmarking graph-based analysis and prediction tasks by converting relational database data into graphs.
Our code is available at GitHub.
Overview
RDB2G-Bench provides comprehensive performance evaluation data for graph neural network models applied to relational database tasks. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/kaistdata/RDB2G-Bench.self-self-distillation
self-self-distillation
Per-question teacher/student reward-delta annotations for verifier-free self-self-distillation,
computed on the sky_work_math subset of
PrimeIntellect/SYNTHETIC-2-RL with
Qwen/Qwen3-4B.
For each problem we draw k=8 rollouts in thinking-on (teacher) and thinking-off (student) modes at
identical sampling (temperature 0.7 / top_p 0.8), grade each against the ground truth, and record the
per-mode expected reward and their difference (delta = R_teacher -… See the full description on the dataset page: https://huggingface.co/datasets/rdavion/self-self-distillation.RDB_PFN
RDB_PFN Datasets
This repository contains the synthetic pre-training and benchmark datasets for RDB-PFN, as presented in the paper Relational In-Context Learning via Synthetic Pre-training with Structural Prior.
Paper: https://arxiv.org/abs/2603.03805
GitHub Repository: https://github.com/MuLabPKU/RDBPFN
Description
RDB-PFN is the first relational foundation model trained purely via synthetic data. Because high-quality relational databases (RDBs) are often private… See the full description on the dataset page: https://huggingface.co/datasets/yamboo/RDB_PFN.acc_rd_s1-gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.chembl-2025-randomized-smiles-cleaned-rdkit-descriptorspiper_3apples_rdm_positionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "piper_6dof",
"total_episodes": 400,
"total_frames": 117842,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:400"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Faless/piper_3apples_rdm_position.gim_arm_industrial_assembly_teleop_3cam_joint
gim_arm_industrial_assembly_teleop_3cam_joint
Human teleoperation demonstrations of a single 6-DOF GIM XL arm with a parallel
gripper, recorded at 30 Hz with three RGB cameras.
Task prompt: assemble the industrial parts
Episodes
272
Frames
147,066
Duration
~1.4 h
Control / logging rate
30 Hz
Episode length
301-1359 frames (median 484, mean 541)
Cameras
3 x Intel RealSense D405, 720x1280 colour
Format
LeRobot v3.0
Robot
GIM XL 6-DOF, right arm… See the full description on the dataset page: https://huggingface.co/datasets/RDLwicked/gim_arm_industrial_assembly_teleop_3cam_joint.bokeh-eval-rd-kfix-fullbokeh-eq4-rd-nossopiper_3apples_rdm_position_2camsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "piper_6dof",
"total_episodes": 400,
"total_frames": 117842,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:400"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Faless/piper_3apples_rdm_position_2cams.gim_arm_socket_plug_teleop_3cam_joint
gim_arm_socket_plug_teleop_3cam_joint
Human teleoperation demonstrations of a single 6-DOF GIM XL arm with a parallel
gripper, recorded at 30 Hz with three RGB cameras.
Task prompt: pick up the plug and insert it into the socket
Episodes
376
Frames
188,393
Duration
~1.7 h
Control / logging rate
30 Hz
Episode length
273-1136 frames (median 444, mean 501)
Cameras
3 x Intel RealSense D405, 720x1280 colour
Format
LeRobot v3.0
Robot
GIM XL 6-DOF, right… See the full description on the dataset page: https://huggingface.co/datasets/RDLwicked/gim_arm_socket_plug_teleop_3cam_joint.bokeh-eval-rd-rotac-fullbokeh-eq4-rd-kfixbokeh-eq4-rd-oficialbokeh-eq4-rd-semtreinohalf-of-chembl-2025-randomized-smiles-cleaned-rdkit-descriptorsagenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/rdubwiley/agenda-parser-tool-traces.bokeh-eq4-rd-rotac60krde-72-dataset
RDE-72: Rotating Detonation Engine Spatiotemporal Dataset
Overview
RDE-72 is the first open dataset of full 2D spatial fields from Rotating Detonation Engine (RDE) CFD simulations. It combines 12 high-fidelity OpenFOAM reactingFoam cases with 60 synthetic cases generated via a Conditional Variational Autoencoder (CVAE), enabling neural surrogate modeling and operating envelope mapping.
Key feature: Unlike prior RDE datasets that provide only bulk statistics (mean pressure… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/rde-72-dataset.gim_arm_pin_insertion_teleop_3cam_joint
gim_arm_pin_insertion_teleop_3cam_joint
Human teleoperation demonstrations of a single 6-DOF GIM XL arm with a parallel
gripper, recorded at 30 Hz with three RGB cameras.
Task prompt: pick up the pin and insert it into the hole
Episodes
175
Frames
159,347
Duration
~1.5 h
Control / logging rate
30 Hz
Episode length
398-2664 frames (median 758, mean 911)
Cameras
3 x Intel RealSense D405, 720x1280 colour
Format
LeRobot v3.0
Robot
GIM XL 6-DOF, right arm… See the full description on the dataset page: https://huggingface.co/datasets/RDLwicked/gim_arm_pin_insertion_teleop_3cam_joint.libero_test_lerobot_pathmask_rdpThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1434,
"total_frames": 208483,
"total_tasks": 33,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1434"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jesbu1/libero_test_lerobot_pathmask_rdp.gim_arm_pick_n_place_teleop_3cam_joint
gim_arm_pick_n_place_teleop_3cam_joint
Human teleoperation demonstrations of a single 6-DOF GIM XL arm with a parallel
gripper, recorded at 30 Hz with three RGB cameras.
Task prompt: pick up the object and place it into the box
Episodes
135
Frames
123,376
Duration
~1.1 h
Control / logging rate
30 Hz
Episode length
388-2971 frames (median 578, mean 914)
Cameras
3 x Intel RealSense D405, 720x1280 colour
Format
LeRobot v3.0
Robot
GIM XL 6-DOF, right arm… See the full description on the dataset page: https://huggingface.co/datasets/RDLwicked/gim_arm_pick_n_place_teleop_3cam_joint.bokeh-eval-rd-semtreino-fullbokeh-eq4-rd-nofilterbokeh-eval-rd-oficial-full
