datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rocketleague-analysis
Rocket League Analysis
Local Rocket League replay analysis using Ballchasing API exports and plain DuckDB.
The report is meant to answer one practical question: what should I work on next from my saved replay sample?
Quick Start
uv sync --locked
UV_CACHE_DIR=/tmp/rocketleague-uv-cache \
uv run --locked pytest -v
uv run --locked python scripts/analyze_scenarios.py \
--replay-dir /path/to/Rocket\ League/TAGame/Demos \
--limit 10
Start with CONTRIBUTING.md… See the full description on the dataset page: https://huggingface.co/datasets/edmundmiller/rocketleague-analysis.EDM-DOCK-training-minimizedimagenet_edm2SWE-V-SIFzinc250kzinc250k contains the 250k molecule subset used in the "Automatic Chemical Design Using a
Data-Driven Continuous Representation of Molecules" paper (doi:10.1021/acscentsci.7b00572).
The dataset contains the original columns from
https://github.com/aspuru-guzik-group/chemical_vae/blob/main/models/zinc_properties/250k_rndm_zinc_drugs_clean_3.csv,
namely smiles, logP, QED, and SAS and an additional selfies column.
This dataset can be used for benchmarking chemical Language Models or training… See the full description on the dataset page: https://huggingface.co/datasets/edmanft/zinc250k.waymo-ipace-detector-dataset
Waymo I-PACE Vehicle Detection Dataset
YOLO-format object detection dataset for detecting Waymo autonomous vehicles (Jaguar I-PACE) in Austin traffic camera images.
Dataset Structure
├── images/
│ ├── train/ # Training images (JPG)
│ └── val/ # Validation images (JPG)
├── labels/
│ ├── train/ # YOLO format annotations (TXT)
│ └── val/ # YOLO format annotations (TXT)
└── dataset.yaml # YOLO configuration
Label Format
YOLO… See the full description on the dataset page: https://huggingface.co/datasets/EDM25/waymo-ipace-detector-dataset.shapes3d_x0pred_edm_sigmaCausal-Reasoning-Bench_CRBench
🦙 Causal Reasoning Benchmark (CRBench)
CRBench is a benchmark for evaluating process-level causal failures in
Chain-of-Thought (CoT) reasoning.
Rather than treating incorrect reasoning traces as homogeneous failures,
CRBench characterizes erroneous dependencies among intermediate reasoning
steps through a step-level causal-error taxonomy. It is designed to evaluate
whether reasoning methods can identify and correct structured causal failures
that arise during the reasoning… See the full description on the dataset page: https://huggingface.co/datasets/EdmondFU/Causal-Reasoning-Bench_CRBench.hyvoxpopuli
HyVoxPopuli
HyVoxPopuli is an open Armenian speech dataset (~6 hours, 16 kHz mono) with raw and normalized transcripts. It targets ASR and TTS research for Eastern Armenian (hy-AM).
Note: Despite the name, this release is not the official Facebook VoxPopuli parliament corpus. Audio is literary narration (two voice actors) segmented into short clips. Update citations and experiments accordingly.
Dataset summary
Rows
623
Train / Val / Test
498 / 62… See the full description on the dataset page: https://huggingface.co/datasets/Edmon02/hyvoxpopuli.edmunds_cube_pickedm-cuefrom datasets import load_dataset
captions = load_dataset("disco-eth/edm-cue")
What is EDM-CUE?
The EDM-CUE dataset contains metadata for ~5k EDM tracks. Cue points are essential for DJs, so we asked the question "can they be placed by a learned system?" To Answer this question we gathered 21k cue points manually placed by human experts, and provide them in this dataset for future use.
To cite this dataset or for more information, please see Cue Point Estimation using Object… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/edm-cue.edmt-dictionary-english-khmer-llm-translated
EDMT English-Khmer Dictionary Dataset (LLM-Translated via Gemini 3.7 & Quality Evaluated)
A comprehensive, high-coverage English-to-Khmer bilingual dictionary dataset containing 176,064 entries and 110,504 distinct English headwords, built upon the open-source EDMT Dictionary Database (Webster's Revised Unabridged Dictionary).
Every word definition and translation has been translated and adapted into natural, grammatically sound Khmer using Gemini 3.7 Flash. In addition, both… See the full description on the dataset page: https://huggingface.co/datasets/namae101/edmt-dictionary-english-khmer-llm-translated.edmunds_cube_pick_goalRTPSpeechPortuguese Speech
This dataset aims to provide people with European Portuguese audio and textual data pairs, which can be used to fine-tune large language models.
These Portuguese recordings are from RTP (Rádio e Televisão de Portugal), which we have broken down and transcribed into short sentences.
Citation
If you use this dataset, please cite:
L. M. Hoi, Y. Sun and S. K. Im, "An Automatic Speech Segmentation Algorithm of Portuguese based on Spectrogram Windowing," 2022 IEEE World AI IoT… See the full description on the dataset page: https://huggingface.co/datasets/edmond5995/RTPSpeech.eval_opentrajectorydit151_unseenThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 4220,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/eval_opentrajectorydit151_unseen.shapes3d-small-x0-ideal-edmeval_opentrajectorydit151_cubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 4998,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/eval_opentrajectorydit151_cube.eval_opentrajectorydit151_cupThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 4401,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/eval_opentrajectorydit151_cup.edmunds-car-ratingseval_opentrajectorydit151_toycarThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 5537,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/eval_opentrajectorydit151_toycar.robotics-papers-vecdb
Robotics Papers Vector Database
Semantic search over 63,381 academic papers from 30 conference-year combinations in robotics, CV, and ML.
Contents
Embeddings: BAAI/bge-m3 (1024 dimensions) via SiliconFlow
Conferences: CoRL, CVPR, ECCV, ICCV, ICLR, ICML, ICRA, IROS, NeurIPS, RSS, WACV (2023-2026)
Fields: title, abstract, author, conference, year, arxiv, github, citations, keywords
Format: LanceDB (compacted, single fragment)
Usage… See the full description on the dataset page: https://huggingface.co/datasets/EdmondValar/robotics-papers-vecdb.dmd_cifar10_edm_distillation_datasetmultilingual-task-oriented-dialog
Multilingual Task-Oriented Dialog Data
Directory structure
This dataset consists of 3 directories:
en contains the English data
es contains the Spanish data
th contains the Thai data
In each directory, you'll find a file for each of the train/dev/test splits as used in our paper.
File format
PYTEXT parquet FORMAT
Each parquet file contains following 5 columns: intent label, the slot annotations in a comma-separated list with the format <start token>:<end… See the full description on the dataset page: https://huggingface.co/datasets/edmdias/multilingual-task-oriented-dialog.Emotional_MusicEmotional Music MIDI Samples
This dataset provides test samples of emotional music generated from original MIDI files. The generative algorithm can be found in the following research paper. This algorithm is able to generate six emotions (anger, fear, happiness, sadness, surprise, and tenderness) at a time.
Citation
If you use this dataset, please cite:
Pinyan Li and Lap Man Hoi and Yapeng Wang and Xu Yang and Yuqi Sun (2026). Transforming melodies into natural language lyrics and emotional… See the full description on the dataset page: https://huggingface.co/datasets/edmond5995/Emotional_Music.so100_block_in_cup-trajectoryThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 21782,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/so100_block_in_cup-trajectory.so100_battery-trajectoryThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 24263,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/so100_battery-trajectory.toycars_trialThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 11,
"total_frames": 7898,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:11"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/toycars_trial.so100_pickplace_1camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 22024,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/so100_pickplace_1cam.eval_ACT-Trajectory-PickPlace_v1_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 8,
"total_frames": 3861,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/eval_ACT-Trajectory-PickPlace_v1_2.so100_stack_green_cube-trajectoryThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100",
"total_episodes": 15,
"total_frames": 6619,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/so100_stack_green_cube-trajectory.
