datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recam-lerobotcelebA_spoof
Dataset Card for "celebA_spoof"
More Information needed
SportsSlomo-CVS
🎥 SportsSloMo-CVS Dataset
This repository contains the dataset presented in the paper Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor.
The Complementary Vision Sensor (CVS), known as Tianmouc, captures synchronized RGB frames together with high-frame-rate, multi-bit spatial difference (SD, encoding structural edges) and temporal difference (TD, encoding motion cues) data within a single RGB exposure. This dataset facilitates research in RGB… See the full description on the dataset page: https://huggingface.co/datasets/mypThu/SportsSlomo-CVS.uva_spoj_rawduplexgen-spoken
DuplexGen Spoken
Rendered spoken audio for the DuplexGen turn-taking dialogues — the exact
set of clips used to fine-tune the full-duplex model (PP-DG) in
DuplexGen: Adaptive Synthesis of Human–AI Turn-Taking Dialogues.
Each clip is a full render of one generated dialogue variation: the mixed
two-speaker dialogue audio, the isolated per-turn utterances, and the inserted
backchannel clips, plus per-clip metadata. Audio is synthesized with
Chatterbox TTS; the dialogue
text it… See the full description on the dataset page: https://huggingface.co/datasets/DuplexGen/duplexgen-spoken.spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!
Column… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/spotify-tracks-dataset.sports-calendar
Sports Calendar Feeds
Public calendar and scoreboard artifacts for MLB, NHL, NFL, College Football
(NCAA Division I), the English Premier League, and selected men's cricket
competitions.
The feeds update on a deterministic four-hour GitHub Actions schedule. Calendar
events are transparent (Free), use stable first-party UIDs, update completed
games with final scores, and mark called-off fixtures as cancelled.
For normal calendar use, choose the Daily Scoreboard feed. It condenses… See the full description on the dataset page: https://huggingface.co/datasets/karunapu/sports-calendar.govdocs1-pdf-source
govdocs1: source PDF files
[!NOTE]
Converted versions of other document types (word, txt, etc) are available in this repo
This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd.
Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details
5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.sports-trends-dataset
⚽🏀🎾🏏 Sports-Trends Dataset
A leakage-safe, multi-sport match data lake — raw fixtures → engineered features → training splits.
The data backbone of Ruslan Magana Sports Intelligence — refreshed automatically every day.
TL;DR — A continuously-updated, medallion-architecture data lake for football,
basketball, tennis and cricket: immutable raw ingests, cleaned/standardized layers, an
engineered feature store, and ready-to-train chronological splits in… See the full description on the dataset page: https://huggingface.co/datasets/ruslanmv/sports-trends-dataset.sportsmot
SportsMOT
SportsMOT is a large-scale dataset for single-camera multi-object tracking (MOT) in sports videos. It focuses on tracking players in professional sports scenes, where targets exhibit fast and variable-speed motion, frequent occlusion, motion blur, camera motion, and similar team uniforms.
This Hugging Face repository provides SportsMOT in a MOTChallenge-style structure for research, benchmarking, training, and evaluation of multi-object tracking systems in sports… See the full description on the dataset page: https://huggingface.co/datasets/Lekim89/sportsmot.btcusdt_spot_1m_03_2023_to_12_2025ml_spoken_wordsMultilingual Spoken Words Corpus is a large and growing audio dataset of spoken
words in 50 languages collectively spoken by over 5 billion people, for academic
research and commercial applications in keyword spotting and spoken term search,
licensed under CC-BY 4.0. The dataset contains more than 340,000 keywords,
totaling 23.4 million 1-second spoken examples (over 6,000 hours). The dataset
has many use cases, ranging from voice-enabled consumer devices to call center
automation. This dataset is generated by applying forced alignment on crowd-sourced sentence-level
audio to produce per-word timing estimates for extraction.
All alignments are included in the dataset.physical-ai-bench-step-30000-videos
PhysicalAIBench Step 30000 Model Comparison
This repository contains two filename-aligned sets of 5,220 MP4 outputs from
step_30000 evaluation runs. The files are presented through the
companion PhysicalAI Video Gallery.
Model sets
Directory
Model
Files
Bytes
videos/
DC-AE v0.2, Cosmos encoder + causal decoder
5,220
4,244,965,541
videos_wan22_vae/
Wan 2.2 VAE, phase 3 PDX
5,220
4,818,900,749
The two directories have an exact 1:1 basename match.… See the full description on the dataset page: https://huggingface.co/datasets/spongy/physical-ai-bench-step-30000-videos.tempspoc_rawkitchen_rack_combo_v2_spoon_onlyThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 183,
"total_frames": 75265,
"total_tasks": 1,
"total_videos": 549,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:183"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/kitchen_rack_combo_v2_spoon_only.wikipedia-20230901.en-deduped
wikipedia - 20230901.en - deduped
purpose: train with less data while maintaining (most) of the quality
This is really more of a "high quality diverse sample" rather than "we are trying to remove literal duplicate documents". Source dataset: graelo/wikipedia.
configs
default
command:
python -m text_dedup.minhash \
--path $ds_name \
--name $dataset_config \
--split $data_split \
--cache_dir "./cache" \
--output $out_dir \
--column $text_column \… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/wikipedia-20230901.en-deduped.cad-bench-freecad-agent-initial-runsAI2_Alphabot_2_stir_spoon
AI2_Alphabot_2_stir_spoon
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 405
Total Frames: 334844
FPS: 30
Dataset Size: 14.30 GB
Robot Name: AI2_Alphabot_2
End-Effector Type: two_finger_end_effector
Teleoperation Type: vr_controller
Sensors: cam_front_chest_rgb,
cam_front_head_rgb,
cam_left_wrist_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_stir_spoon.code_contests_instruct
Dataset Card for "code_contests_instruct"
The deepmind/code_contests dataset formatted as markdown-instruct for text generation training.
There are several different configs. Look at them. Comments:
flesch_reading_ease is computed on the description col via textstat
hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater
min-cols drops all cols except language and text
possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.spooky-author-identificationLong-Data-Col-rp_pile_pretrain
Dataset Card for "Long-Data-Col-rp_pile_pretrain"
This dataset is a subset of togethercomputer/Long-Data-Collections, namely the rp_sub.jsonl.zst and pile_sub.jsonl.zst files from the pretrain split.
Like the source dataset, we do not attempt to modify/change licenses of underlying data. Refer to the source dataset (and its source datasets) for details.
changes
as this is supposed to be a "long text dataset", we drop all rows where text contains <= 250 characters.… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/Long-Data-Col-rp_pile_pretrain.face-anti-spoofing-dataset
Face Antispoofing dataset for liveness detection
Anti-Spoofing dataset: live, replay, cut, print, 3D masks - large-scale face anti spoofing
This dataset delivers a single, end-to-end resource for training and benchmarking facial liveness-detection systems. By aggregating live sessions and eleven realistic presentation-attack classes into one collection, it accelerates development toward iBeta Level 1/2 compliance and strengthens model robustness against the full spectrum of spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/face-anti-spoofing-dataset.SpokenWOZ-Test-Audiommu_tess_spoc
mmu_tess_spoc HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_tess_spoc.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_tess_spoc.SEA-Spoof
SEA-Spoof
Access And License
Access requires author approval. Please email the authors before requesting or using the dataset:
wu_jinyang@a-star.edu.sg
imcc.sg@gmail.com
This dataset is released for non-commercial academic research only.
Use is restricted to academic institutions and approved research users. Commercial use is not permitted, and this dataset may not be used by commercial companies or for commercial products, services, model training, evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Jack-ppkdczgx/SEA-Spoof.spotify-songsconsumer-finance-complaints
BEE-spoke-data/consumer-finance-complaints
consumer-finance-complaints but in a format that actually works.
Pulled Feb 2024
binance-top50-spot-v1
Binance Top 50 Backtesting Dataset
Built at: 2026-05-21T11:53:45.468482+00:00
Parameters
Lookback: 1 days
Top N: 3
Trade Types: spot, um
Data Types: klines, aggTrades
Build Status
SPOT: 3 symbols
klines: 3/3 healthy
aggTrades: 3/3 healthy
UM: 3 symbols
klines: 3/3 healthy
aggTrades: 3/3 healthy
fundingRate: 3/3 healthy
SportsTime
SportsTime
SportsTime is a long-form sports video question answering benchmark for temporal compositional reasoning, accepted to ECCV 2026.
It contains 14,326 open-ended QA pairs with 50,000+ step-wise temporal evidence annotations across 1,575 videos and five team sports: basketball, American football, ice hockey, soccer, and volleyball.
Dataset
This Hugging Face dataset repository provides the annotation files, official train/test split, and video files.… See the full description on the dataset page: https://huggingface.co/datasets/Ustiniansy/SportsTime.
