datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
abc_130k_v3_trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"left_arm_joint_1",
"left_arm_joint_2",
"left_arm_joint_3",
"left_arm_joint_4",
"left_arm_joint_5"… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/abc_130k_v3_train.ABC-Pretraining-Data
ABC Pretraining Data
This dataset contains the pretraining data for ABC, an open-source multimodal embedding model that uses a vision-language model backbone to deeply integrate image features with natural language instructions, advancing the state of visual embeddings with natural language control.
This dataset is derived from Google's Conceptual Captions dataset.
Each item in the dataset contains a URL where the corresponding image can be downloaded and mined negatives for… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ABC-Pretraining-Data.CALVIN_ABC_tarABC-1Mabc-testing
ABC Testing dataset
A: A0-2**24
B: B0-2**20
C: C0-2**16
Jee-Chemistry-dataset-with-COTabc_130k_trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.top": {
"dtype": "video",
"shape": [
224,
224,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/typoverflow/abc_130k_train.abcde
ABCDE v1.1 — processed-feature release
Version 1.1: public processed-feature release. Source lexicons are excluded.
This release contains 264,210,076 enriched text records across 31 source tables, plus three unchanged demographic-user tables. It extends the pinned ABCDE snapshot d6797a1c0c935de66c6a022e3d0ed1ea705c4c10.
Applies the committed whole-word HasBPM fix and recalculates the five possessive BPM columns.
Adds 48 moral features per text field across Care, Loyalty, Purity… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/abcde.midi-audio-abc_300smidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5-300s, several subsets with smaller duration
60s
30s
10s)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)
Citation
@misc{jiang2025advancingfoundationmodelmusic… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_300s.abc_130k_v3_smokeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"left_arm_joint_1",
"left_arm_joint_2",
"left_arm_joint_3",
"left_arm_joint_4",
"left_arm_joint_5"… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/abc_130k_v3_smoke.open-web-math
Keiran Paster*, Marco Dos Santos*, Zhangir Azerbayev, Jimmy Ba
GitHub | ArXiv
| PDF
OpenWebMath is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of 6.3 million documents containing a total of 14.7B tokens. OpenWebMath is intended for use in pretraining and finetuninglarge language models.
You can download the dataset using Hugging Face:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ABC321000/open-web-math.midi-audio-abc_60smidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5-60s, sampled from the full set with max 300s duration)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)
Citation
@misc{jiang2025advancingfoundationmodelmusic… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_60s.CALVIN-3D_PCD-ABC_D
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABC_D.midi-audio-abc_longmidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5 min - 2 hours, less than 5 min data are in 300s
and there are several subsets with smaller duration
60s
30s
10s)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_long.CALVIN-3D_PCD-ABCD_D
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABCD_D.Hand_Tools
ABC-iRobotics/Hand_Tools: dataset 0000
Metric RGB-D scene dataset generated with the Label Factory workflow.
The files for this capture are stored below the 0000/ directory so
multiple numbered datasets can coexist in this repository.
Training configurations
depth_estimation: RGB input, metric depth target, intrinsics and depth units
instance_segmentation: RGB input, instance-mask target, boxes and annotations
object_pose_estimation: RGB-D input, masks, camera… See the full description on the dataset page: https://huggingface.co/datasets/ABC-iRobotics/Hand_Tools.therapy-conversations-multiturn
Combined Dr. AURA Therapy Conversations Dataset
This dataset is 100% AI-generated for research and educational purposes only. It is not intended to provide medical, psychological, or therapeutic advice. Always consult a qualified healthcare professional or doctor for any mental health concerns or medical issues. AI-generated content may contain errors or inaccuracies.
warning ⚠️: the LENGTH of each conversation MAY VARY (eg. 7 or 8 or 9 or 10 etc. turns in each row). And SOME END… See the full description on the dataset page: https://huggingface.co/datasets/Abc7347/therapy-conversations-multiturn.abc_subset_rbmw_abcyuth_324VIGIA-QA-datasetr_abcyuth_324abcABC-Bench
ABC-Bench
💻 Code |
📑 Paper |
📝 Blog
📖 Overview
ABC-Bench is a benchmark for Agentic Backend Coding. It evaluates whether code agents can explore real repositories, edit code, configure environments, deploy containerized services, and pass external end-to-end API tests (HTTP-based integration tests) across realistic backend stacks.
📊 Benchmark Composition
🚀 Why ABC-Bench?
End-to-End Lifecycle: repository exploration → code… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/ABC-Bench.midi-audio-abc_30smidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5-30s, sampled from the full set with max 300s duration)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)
Citation
@misc{jiang2025advancingfoundationmodelmusic… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_30s.gaokao-sft-chinese-strict-abcd-v3
Gaokao SFT Chinese Strict ABCD V3
This dataset is the cleaned Chinese SFT release that keeps only single-choice samples where A, B, C, and D all have explicit option-level analysis.
Composition
Total samples: 88466
Train samples: 86670
Validation samples: 1796
Subject Counts
{
"biology": 33104,
"chemistry": 35796,
"english": 174,
"general_exam": 7982,
"geography": 888,
"history": 229,
"physics": 9986,
"politics": 307
}
Fields
id… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd-v3.ABC-K562abcplaneThis is a dataset card for the ABC-plane dataset used in PaCo and UniCo for structured shape completion.
The dataset is provided in preprocessed form for training and evaluation with the released code.
Citations
If you use this dataset in scientific work, please consider citing the papers:
@article{chen2026unico,
title={Unified Primitive Proxies for Structured Shape Completion},
author={Zhaiyu Chen and Yuqing Wang and Xiao Xiang Zhu},
journal={arXiv preprint arXiv:2601.00759}… See the full description on the dataset page: https://huggingface.co/datasets/chenzhaiyu/abcplane.CALVIN_ABCabcd567_ziprepro-abc-bench-an-agentic-bio-capabilities-benchmark-for-biosecurity-traces
Agent traces
Agent sessions published from a Trackio Logbook.
