datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!
Column… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/spotify-tracks-dataset.govdocs1-pdf-source
govdocs1: source PDF files
[!NOTE]
Converted versions of other document types (word, txt, etc) are available in this repo
This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd.
Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details
5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.sports-trends-dataset
⚽🏀🎾🏏 Sports-Trends Dataset
A leakage-safe, multi-sport match data lake — raw fixtures → engineered features → training splits.
The data backbone of Ruslan Magana Sports Intelligence — refreshed automatically every day.
TL;DR — A continuously-updated, medallion-architecture data lake for football,
basketball, tennis and cricket: immutable raw ingests, cleaned/standardized layers, an
engineered feature store, and ready-to-train chronological splits in… See the full description on the dataset page: https://huggingface.co/datasets/ruslanmv/sports-trends-dataset.btcusdt_spot_1m_03_2023_to_12_2025kitchen_rack_combo_v2_spoon_onlyThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 183,
"total_frames": 75265,
"total_tasks": 1,
"total_videos": 549,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:183"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/kitchen_rack_combo_v2_spoon_only.code_contests_instruct
Dataset Card for "code_contests_instruct"
The deepmind/code_contests dataset formatted as markdown-instruct for text generation training.
There are several different configs. Look at them. Comments:
flesch_reading_ease is computed on the description col via textstat
hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater
min-cols drops all cols except language and text
possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.mmu_tess_spoc
mmu_tess_spoc HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_tess_spoc.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_tess_spoc.SP_OrderingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 24037,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/SP_Ordering.AIRBOT_MMK2_beauty_sponge_and_cake_to_place
AIRBOT_MMK2_beauty_sponge_and_cake_to_place
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_beauty_sponge_and_cake_to_place.AIRBOT_MMK2_storage_for_building_blocks_and_beauty_sponges
AIRBOT_MMK2_storage_for_building_blocks_and_beauty_sponges
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_for_building_blocks_and_beauty_sponges.spot-terrain-dataset
Spot Dataset
nangang_sports_centerspotify-tracks-lite
Context
This dataset consists of 24000 tracks from 30 genres, and is a shrunk version of maharshipandya/spotify-tracks-dataset dataset. All non-heuristic data is cut and cleaned for better usability and performance.
All data taken from Spotify API and is open source.
This dataset can be used to train prediction models based on user preferences, or categorise tracks by corresponding heuristic.
Column Description
danceability: Danceability describes how suitable a track is… See the full description on the dataset page: https://huggingface.co/datasets/engels/spotify-tracks-lite.spotify_audio_features
Spotify Tracks & Audio Features Dataset
Overview
This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research.
Data Source
The raw data for this dataset was originally gathered and hosted by Anna's Archive.
Original Blog Post:… See the full description on the dataset page: https://huggingface.co/datasets/ozefe/spotify_audio_features.sportsbookish-daily-odds
SportsBookISH Daily Kalshi vs Sportsbook Odds
Real-time pricing snapshot comparing Kalshi event-contract probabilities against US sportsbook consensus across nine sports.
Description
Daily-refreshed JSON / CSV export of every active Kalshi market alongside the de-vigged book median across 13+ US sportsbooks. Covers golf (PGA Tour), NFL, NBA, MLB, NHL, EPL, MLS, UEFA Champions League, and FIFA World Cup.
Source
Live data plane:
JSON:… See the full description on the dataset page: https://huggingface.co/datasets/kennyhyder/sportsbookish-daily-odds.gnss-jamming-spoofing-detection
GNSS Jamming & Spoofing Detection Dataset
A physics-informed synthetic dataset for detecting GPS/GNSS cyber-attacks
(Jamming and Spoofing) from satellite-signal features. Built for the
GNSS Guardian project — Introduction to Data Science final project.
Overview
14,850 samples across 450 scenarios × 33 time-steps each
3 balanced classes: Normal / Jamming / Spoofing (4,950 each)
26 columns: multi-constellation signal features + attack metadata + text descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Omrilevi123/gnss-jamming-spoofing-detection.AIRBOT_MMK2_place_the_sponge_and_wet_wipes
AIRBOT_MMK2_place_the_sponge_and_wet_wipes
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_sponge_and_wet_wipes.top-hits-spotifyAIRBOT_MMK2_storage_box_for_mouse_and_sponge
AIRBOT_MMK2_storage_box_for_mouse_and_sponge
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_box_for_mouse_and_sponge.spotify-huge-track-analysis-dataset
Spotify Track Analysis Dataset
General Description
This dataset provides a large-scale, research-oriented analytical representation of Spotify music data.
It is centered on tracks as musical recordings (track_id), while preserving explicit artist attribution as defined by Spotify’s native credit model.
Each row corresponds to a track–artist association, identified by:
a Spotify track identifier (track_id)
a credited artist name (artist_name)
A single track may appear on… See the full description on the dataset page: https://huggingface.co/datasets/GildasLeDrogoff/spotify-huge-track-analysis-dataset.spotify_audio_features_partitionedethusdt_spot_1m_05_2021_to_03_2026
ETHUSDT Spot 1-Minute OHLCV (May 2021 - Mar 2026)
Overview
1-minute OHLCV candlestick data for the ETH/USDT spot pair on Binance, covering May 1, 2021 to February 28, 2026.
Rows: 2,541,600
Completeness: 100.00%
Sources
Period
Source
Notes
Full dataset
Binance Data Collection
Monthly kline ZIPs
2021-08-13 02:00-06:29
Bybit API
270 bars filled from Bybit ETHUSDT spot (Binance maintenance)
2021-09-29 07:00-08:59
Bybit API
120 bars filled from… See the full description on the dataset page: https://huggingface.co/datasets/Torch-Trade/ethusdt_spot_1m_05_2021_to_03_2026.sponsorblock-youtube-metadata-2024
SponsorBlock YouTube Metadata Dataset
A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos.
Contains the top videos from the SponsorBlock database that had data added in the year 2024.
Quick Stats
Metric
Value
Total videos
154,536
Videos with subtitles
62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.R1_Lite_move_the_position_of_the_spoon
R1_Lite_move_the_position_of_the_spoon
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
place
pick
grasp
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_move_the_position_of_the_spoon.SpotifyDataavs-spot
Dataset Card for AVS-Spot Benchmark
This dataset is associated with the paper: "Understanding Co-Speech Gestures in-the-wild"
📝 ArXiv: https://arxiv.org/abs/2503.22668
🌐 Project page: https://www.robots.ox.ac.uk/~vgg/research/jegal
💻 Code: https://github.com/Sindhu-Hegde/jegal
We present JEGAL, a Joint Embedding space for Gestures, Audio and Language. Our semantic gesture representations can be used to perform multiple downstream tasks such as cross-modal retrieval… See the full description on the dataset page: https://huggingface.co/datasets/sindhuhegde/avs-spot.record-test-2cam-pink-spongesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 16502,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YoshitoMori/record-test-2cam-pink-sponges.Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.spotify-tracks-datasetsponge_marker_merged_videoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so_follower",
"total_episodes": 100,
"total_frames": 78522,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mkpongm/sponge_marker_merged_video.
