datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pointodysseyFeb 21: updated to v1.2. Please see https://github.com/y-zheng18/point_odyssey/tree/main for release notes.
aharensanwahakarenaiseason2
Bangumi Image Base of Aharen-san Wa Hakarenai Season 2
This is the image base of bangumi Aharen-san wa Hakarenai Season 2, we detected 65 characters, 6341 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/aharensanwahakarenaiseason2.csvfiverr-gigs
Fiverr Gigs — community dataset
Public Fiverr listing metadata collected via the
fiverr-gig-optimizer
Claude Code skill. Grown via opt-in contributions from users who run
contribute.py. Licensed CC-BY-4.0.
What's in it
Each record follows the canonical schema (title, category/subcategory, tier
prices, delivery days, tags, rating, review_count, gig_count_in_search,
currency). Prices are normalized to USD.
Privacy
Contributions are anonymized before… See the full description on the dataset page: https://huggingface.co/datasets/Ahad690/fiverr-gigs.growthkit-trends
GrowthKit Trends
A community, opt-in, federated dataset of public, anonymized short-form
short-form-video trend and benchmark observations, contributed by users of the
open-source GrowthKit Claude
Code skill. It improves GrowthKit's default benchmarks over time so every
founder starts from better, source-tagged ranges instead of fabricated numbers.
Honesty first. GrowthKit never lets a model invent a market metric. Numbers
come from deterministic scripts run on a founder's own… See the full description on the dataset page: https://huggingface.co/datasets/Ahad690/growthkit-trends.rvl_cdipThe RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset consists of 400,000 grayscale images in 16 classes, with 25,000 images per class. There are 320,000 training images, 40,000 validation images, and 40,000 test images.fpp-ml-bench
FPP-ML-Bench: Fringe Projection Profilometry Benchmarking Dataset
The first open-source, photorealistic synthetic dataset for single-shot fringe projection profilometry (FPP), generated using VIRTUS-FPP in NVIDIA Isaac Sim. This dataset enables standardized benchmarking and systematic comparison of deep learning approaches for single-shot 3D depth reconstruction from fringe patterns.
Dataset Summary
Property
Value
Total fringe images
15,600 (52 per viewpoint… See the full description on the dataset page: https://huggingface.co/datasets/aharoon/fpp-ml-bench.crossed_arm_point_clouds
Crossed Arm Point Clouds Dataset
This dataset contains 3D point cloud data captured from a LiDAR scanner for crossed arm classification in the context of robot magic trick performance.
Overview
This dataset was collected for training and evaluating the Crossed Arm Voxel Network (CAVN) architecture, a deep learning model designed for 3D point cloud classification in human-robot interaction magic performances. The data supports classification of human arm positions during… See the full description on the dataset page: https://huggingface.co/datasets/ahanjaya/crossed_arm_point_clouds.app-rank-anchors
App Rank Anchors
Community-federated public app-store calibration anchors for the
AppScope open app-intelligence
stack.
Each row is a public fact — a segment + rank + observed download flow —
derived from the public Google Play realInstalls delta over a time window
paired with an app's chart rank in that window. Pooling these anchors across
self-hosting contributors lets the Garg–Telang download estimator calibrate
absolute scale (scale_b) per (platform, category, country)… See the full description on the dataset page: https://huggingface.co/datasets/Ahad690/app-rank-anchors.aloha-handover-no-mocap
ALOHA 2-Arm Handover (no mocap markers)
163 successful scripted demonstrations of a bimanual box handover in MuJoCo,
recorded on the ALOHA / ViperX 300 dual-arm platform. Arm A picks a box off the
table and hands it to arm B, which retracts with it.
Why "no mocap"
This is a re-collection of
Ahaskar04/aloha-handover-data
with a visual confound removed.
The scripted expert drives the arms through MuJoCo mocap bodies (left/target,
right/target) — 2 cm spheres that… See the full description on the dataset page: https://huggingface.co/datasets/Ahaskar04/aloha-handover-no-mocap.alltracker_dataThis repo contains the data we produced/postprocessed as part of AllTracker: Efficient Dense Point Tracking at High Resolution.
This data is used by the training scripts in our github repo, and leads to the models in our model page, which you can test in our Gradio demo.
viperx-3arm-handover-demo
COLA 3-Arm Sequential Handover — Scripted Demonstrations
1,500 simulated demonstrations of a sequential 3-arm handover task in
MuJoCo, generated with a hand-tuned waypoint policy. Three ViperX 300s
robot arms (A → B → C) cooperate to lift an object off a table, hand it
between two arms, and place it into a box rigidly attached to the third
arm.
This dataset is intended for multi-agent imitation learning, coordination
research, and as the 3-arm extension of the 2-arm… See the full description on the dataset page: https://huggingface.co/datasets/Ahaskar04/viperx-3arm-handover-demo.aharensanwahakarenai
Bangumi Image Base of Aharen-san Wa Hakarenai
This is the image base of bangumi Aharen-san wa Hakarenai, we detected 43 characters, 5875 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/aharensanwahakarenai.oai-5g-srs-ranging-dataset
OAI 5G NR SRS Ranging Captures
Uplink Sounding Reference Signal (SRS) channel-estimate captures from a monolithic 5G NR
software-defined-radio testbed, collected for SRS-based ranging experiments.
The gNB (OpenAirInterface on a USRP X410, Band n78, 40 MHz / 106 PRB) configures each UE
to transmit SRS; the gNB's per-SRS frequency-domain channel estimate, oversampled IDFT CIR,
and ToA estimate are streamed off the PHY via OAI's T_tracer and recorded at a series of
known… See the full description on the dataset page: https://huggingface.co/datasets/ahancock516/oai-5g-srs-ranging-dataset.AHA-WAM-SO101-HIL-training-assets
AHA-WAM SO101 Plug HIL Training Assets
Reproducibility bundle for the SO101 power-adapter insertion experiments.
It contains a compact, ZIP-based representation of the paths expected by the
AHA-WAM training configuration:
the released AHA-WAM-pretrained.pt initialization checkpoint;
so101_ahawam_plug.zip (the downloader extracts only task 02 and 04);
ahawam_hil_raw.zip (raw rich-v1/v2/v3 HG-DAgger recordings);
the fixed 02+04 action/state normalization statistics;
cached T5… See the full description on the dataset page: https://huggingface.co/datasets/Jill111/AHA-WAM-SO101-HIL-training-assets.cola-handover-demosWan21-1.3B-2000promptaloha-handover-datadeepscaler-verl-aha-momentAHA-MEMES
AHA-Memes
A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
Hateful memes carry their meaning in the interaction between an image and the text
laid over it, and often through cultural references that neither modality states
outright. Arabic has been badly served here: the meme resources that exist
annotate propaganda or coarse "harmful content", not who is being attacked or how.
AHA-Memes is a benchmark of 5,000 Arabic memes, each annotated by trained… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AHA-MEMES.spiritoai-5g-plkg-dataset
OAI 5G PLKG Paired Dataset
PUSCH (uplink) and PDSCH (downlink) channel-estimate captures from an OAI
5G NR SA testbed (USRP X410 gNB + USRP B210 UE, band n78, 106 PRB, PCI=1),
collected for a channel-informed neural PHY secret-key generation (PLKG)
demo paper. Campaigns 1-3: captures are handshake/RRC-setup-phase only
(strict DMRS-comb filtering rejects connected-mode single-layer data on
both plugins). Campaign 4 relaxes this filtering on both plugins and
includes ordinary… See the full description on the dataset page: https://huggingface.co/datasets/ahancock516/oai-5g-plkg-dataset.aha-moment-dataset
verl_deepscaler — DeepScaleR for Verl GRPO Training
Full DeepScaleR training dataset (from agentica-org/DeepScaleR-Preview-Dataset, the most downloaded DeepScaleR dataset with 25.4K downloads), converted to Verl framework format, with the DeepSeek-R1 "aha moment" problem included.
Source
Original: agentica-org/DeepScaleR-Preview-Dataset
Size: 40,315 math problems from AIME, AMC, Omni-MATH, STILL
Included: +1 "aha moment" problem (the exact problem from… See the full description on the dataset page: https://huggingface.co/datasets/safafaf311/aha-moment-dataset.slide2svg
Slide2SVG
Slide2SVG was released as part of the paper “Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling” (AAAI 2026).
Slide2SVG is a real-world dataset for semantic document derendering: transforming a rasterized presentation slide into a structured, editable SVG representation. Curated from publicly available academic conference presentations, it captures diverse design styles, font choices, image placements, and layout configurations found in real… See the full description on the dataset page: https://huggingface.co/datasets/ahazimeh/slide2svg.robocasa_mt4_N216_zarrcountdown-verl-aha-moment
Countdown Dataset for Verl - DeepSeek-R1 "Aha Moment" Reproduction
This dataset is prepared for reproducing DeepSeek-R1's "aha moment" using the Verl framework. It uses the Countdown task from Jiayi-Pan/Countdown-Tasks-3to4.
Overview
The "aha moment" was first observed in DeepSeek-R1-Zero when training on the Countdown task — the model learns to allocate more thinking time and re-evaluate its initial approach rather than rushing to an answer.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sdfffafag3/countdown-verl-aha-moment.cola-pushblock-1033
COLA Push-Block — Two-Agent Coordination (1033 demos)
Teleoperated demonstrations of a two-arm cooperative push-block task in
MuJoCo, collected for the COLA (Coordination via Latent Adapters) paper.
Two 4-DoF arms on opposite ends of a near-frictionless ("icy") table
cooperate to (1) push a small block across the midline and (2) stop it
inside a target zone before it slides off the far edge.
This is the cached, flattened version of the dataset used directly by
the COLA training… See the full description on the dataset page: https://huggingface.co/datasets/Ahaskar04/cola-pushblock-1033.deepscaler-aha-momentsvarah_processedThe input_features are nothing but the values generated after passing the dataset's audio array through a whisper processor's feature extraction and the field 'labels' consists of the tokenized(using whisper tokenizer) ground truths.
The following is the link for what I did with the sarvah dataset and how I trained it on whisper-large-v3-turbo.
The training steps for whisper-large-v3 are same.
https://colab.research.google.com/drive/1oD0v7MWZ9WJqk7tZYThwgTUM85PTEhMN?usp=sharing
AHA-Calvin-1p
