datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RDD2022
RDD2022: Multi-National Road Damage Detection Dataset (4-Class YOLO Export)
Unofficial redistribution of the RDD2022 multi-national road-damage dataset, reduced to the 4-class CRDDC2022 taxonomy and reformatted into a standardized YOLO-compatible directory layout, under the original CC BY-SA 4.0 license.
Disclaimer
This repository is not an official release of the RDD2022 dataset.
RDD2022 was created by Deeksha Arya, Hiroya Maeda, Sanjay Kumar Ghosh… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/RDD2022.rdmap_traj_0swe-marathon
SWE Marathon: Ultra Long-Horizon Software Engineering Tasks
20 ultra long-horizon software-engineering tasks designed to challenge frontier coding agents. Each task ships with a containerized environment, a precise instruction, comprehensive tests, and a reference oracle solution. All tasks pass NOP-baseline / Oracle-fix validation.
Homepage: https://github.com/abundant-ai/swe-marathon
License: Apache 2.0
Format: Harbor task format (task.toml + instruction.md + environment/ +… See the full description on the dataset page: https://huggingface.co/datasets/rdesai2/swe-marathon.UNSW-NB15We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily.
NIDS Datasets
The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/rdpahalavan/UNSW-NB15.lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.1m2r-real-tapeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 56,
"total_frames": 23054,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:56"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-real-tape.sim-handover-corrective-1000-0820This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda-panda",
"total_episodes": 2000,
"total_frames": 303710,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:2000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/sim-handover-corrective-1000-0820.rdkit_featuresRDD2022bhl-impact-gt
FineBooks BHL IMPACT Ground Truth
2,165 page scans from six historical natural-history books, each paired with an expert, ~99.95%-accurate transcription and full page-layout ground truth. A benchmark for OCR, text recognition, and document layout analysis on real historical print.
This dataset is the basis of the BHL OCR Leaderboard, where open OCR models are scored against these transcriptions. As new OCR models are released, they are run through the same evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/rdmpage/bhl-impact-gt.1m2r-tape2-bvThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 56,
"total_frames": 49290,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:56"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-tape2-bv.1m2r-handover2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 44,
"total_frames": 63492,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:44"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-handover2.1m2r-handover2-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "kinova-arx",
"total_episodes": 44,
"total_frames": 63492,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:44"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-handover2-v30.rag-logsff-model-personalitysim-handover-corrective-0819This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda-panda",
"total_episodes": 400,
"total_frames": 60484,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:400"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/sim-handover-corrective-0819.nepali_dataset_llm1m2r-real-diffusion2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 72,
"total_frames": 32350,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:72"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-real-diffusion2.sim-handover-0819This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda-panda",
"total_episodes": 400,
"total_frames": 55966,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:400"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/sim-handover-0819.1m2r-realThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 72,
"total_frames": 32350,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:72"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-real.tartandrive_every1_100pct_rdp
tartandrive_every1_100pct_rdp
Description
Processed tartandrive dataset with filter_every_nth=1, 100% of data, and subsampled with Ramer-Douglas-Peucker algorithm.
Processing Parameters
mateoguaman/tartandrive:
exclude_outliers_pct: 0
filter_by_curvature: false
filter_every_nth: 1
horizon:
100: 0.5
300: 0.5
num_subsampled_points: -1
subsample_method: rdp
tolerance: 15
Dataset Configuration
Train dataset:
mixer:… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/tartandrive_every1_100pct_rdp.1m2r-sweep-prThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5-panda",
"total_episodes": 60,
"total_frames": 18168,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-sweep-pr.RDB2G-Bench
RDB2G-Bench
This is an offical dataset of the paper RDB2G-Bench: A Comprehensive Benchmark for Automatic Graph Modeling of Relational Databases.
RDB2G-Bench is a toolkit for benchmarking graph-based analysis and prediction tasks by converting relational database data into graphs.
Our code is available at GitHub.
Overview
RDB2G-Bench provides comprehensive performance evaluation data for graph neural network models applied to relational database tasks. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/kaistdata/RDB2G-Bench.network-packet-flow-header-payloadEach row contains the information of a network packet and its label. The format is given below:
1m2r-basket-0726-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "kinova-arx",
"total_episodes": 60,
"total_frames": 54027,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-basket-0726-v30.ALARB
ALARB Dataset
ALARB includes a dataset of structured legal cases. Each case lists the facts presented by the plaintiff and defendant, and an explicit step-by-step chain of the argument reasoning of the court leading to a verdict. Cases are linked to individual articles of applicable statutes and regulations.
In our paper, ALARB: An Arabic Legal Argument Reasoning Benchmark, we show how this dataset can be leveraged in a set of legal reasoning tasks.
Cite:… See the full description on the dataset page: https://huggingface.co/datasets/THIQAH-RD/ALARB.Big-Math-RL-Verified-Filteredbritish-museum-rdf-as-csv-2014This repository contains data that was released under the BM license issued in 2014.
self-self-distillation
self-self-distillation
Per-question teacher/student reward-delta annotations for verifier-free self-self-distillation,
computed on the sky_work_math subset of
PrimeIntellect/SYNTHETIC-2-RL with
Qwen/Qwen3-4B.
For each problem we draw k=8 rollouts in thinking-on (teacher) and thinking-off (student) modes at
identical sampling (temperature 0.7 / top_p 0.8), grade each against the ground truth, and record the
per-mode expected reward and their difference (delta = R_teacher -… See the full description on the dataset page: https://huggingface.co/datasets/rdavion/self-self-distillation.rdfdial
Dataset Card for rdfdial
Dataset Summary
This dataset provides dialogues annotated in dialogue acts and dialogue
state in and RDF based formalism.
There is a conversion of sfxdial, dstc2 and multiwoz2.3 datasets
as well as two fully synthetic datasets created from simulated conversations:
camrest-sim and multiwoz-sim.
Original dataset before conversion are available here:
DSTC2: https://github.com/matthen/dstc
Multiwoz 2.3:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/rdfdial.
