datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!
Column… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/spotify-tracks-dataset.btcusdt_spot_1m_03_2023_to_12_2025spot-terrain-dataset
Spot Dataset
spotify_audio_features
Spotify Tracks & Audio Features Dataset
Overview
This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research.
Data Source
The raw data for this dataset was originally gathered and hosted by Anna's Archive.
Original Blog Post:… See the full description on the dataset page: https://huggingface.co/datasets/ozefe/spotify_audio_features.spotify-tracks-lite
Context
This dataset consists of 24000 tracks from 30 genres, and is a shrunk version of maharshipandya/spotify-tracks-dataset dataset. All non-heuristic data is cut and cleaned for better usability and performance.
All data taken from Spotify API and is open source.
This dataset can be used to train prediction models based on user preferences, or categorise tracks by corresponding heuristic.
Column Description
danceability: Danceability describes how suitable a track is… See the full description on the dataset page: https://huggingface.co/datasets/engels/spotify-tracks-lite.top-hits-spotifyspotify-huge-track-analysis-dataset
Spotify Track Analysis Dataset
General Description
This dataset provides a large-scale, research-oriented analytical representation of Spotify music data.
It is centered on tracks as musical recordings (track_id), while preserving explicit artist attribution as defined by Spotify’s native credit model.
Each row corresponds to a track–artist association, identified by:
a Spotify track identifier (track_id)
a credited artist name (artist_name)
A single track may appear on… See the full description on the dataset page: https://huggingface.co/datasets/GildasLeDrogoff/spotify-huge-track-analysis-dataset.spotify_audio_features_partitionedethusdt_spot_1m_05_2021_to_03_2026
ETHUSDT Spot 1-Minute OHLCV (May 2021 - Mar 2026)
Overview
1-minute OHLCV candlestick data for the ETH/USDT spot pair on Binance, covering May 1, 2021 to February 28, 2026.
Rows: 2,541,600
Completeness: 100.00%
Sources
Period
Source
Notes
Full dataset
Binance Data Collection
Monthly kline ZIPs
2021-08-13 02:00-06:29
Bybit API
270 bars filled from Bybit ETHUSDT spot (Binance maintenance)
2021-09-29 07:00-08:59
Bybit API
120 bars filled from… See the full description on the dataset page: https://huggingface.co/datasets/Torch-Trade/ethusdt_spot_1m_05_2021_to_03_2026.avs-spot
Dataset Card for AVS-Spot Benchmark
This dataset is associated with the paper: "Understanding Co-Speech Gestures in-the-wild"
📝 ArXiv: https://arxiv.org/abs/2503.22668
🌐 Project page: https://www.robots.ox.ac.uk/~vgg/research/jegal
💻 Code: https://github.com/Sindhu-Hegde/jegal
We present JEGAL, a Joint Embedding space for Gestures, Audio and Language. Our semantic gesture representations can be used to perform multiple downstream tasks such as cross-modal retrieval… See the full description on the dataset page: https://huggingface.co/datasets/sindhuhegde/avs-spot.SpotifyDatadumpster-diving-spots
Dataset Card for "Dataset of Dumpster Diving Spots"
Community-collected dumpster diving spots and their ratings from https://www.dumpstermap.org (or better https://dumpstermap.herokuapp.com/dumpsters).
Updated monthly using https://github.com/Hitchwiki/dumpster-archiver.
Dataset Details
Dataset Sources
Repository: https://github.com/Debakel/Dumpstermap
Demo: https://www.dumpstermap.org/
Uses
gift and sharing econonmy… See the full description on the dataset page: https://huggingface.co/datasets/Hitchwiki/dumpster-diving-spots.spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
spotify-data-960kembeat_45m_spotify_tracks
Embeat 45M Spotify Tracks
A large-scale music metadata dataset containing 45 million Spotify tracks, combined with the Spotify metadata from Anna's Archive and artist genres from Every Noise at Once.
GitHub project: https://github.com/gdstudio-org/Embeat
Dataset Preview
>>> from datasets import load_from_disk
>>> ds = load_from_disk("GD-Studio/embeat_45m_spotify_tracks")
>>> ds
Dataset({
features: ['track_id', 'track_name', 'isrc', 'popularity', 'explicit'… See the full description on the dataset page: https://huggingface.co/datasets/GD-Studio/embeat_45m_spotify_tracks.Spotify_Audio_features_2.3Mspotify-tracks-datasetspotify_popular_tracksspotify_audio_features
Spotify Tracks & Audio Features Dataset
Overview
This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research.
Data Source
The raw data for this dataset was originally gathered and hosted by Anna's Archive.… See the full description on the dataset page: https://huggingface.co/datasets/0xaf/spotify_audio_features.success_at_spotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 3217,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/FrozenAngel/success_at_spot.qwen3.5-2b-base-blind-spots
Qwen3.5-2B-Base — Blind Spot Analysis (Text + Vision)
Model Tested
Field
Value
Model
Qwen/Qwen3.5-2B-Base
Parameters
2.27 B (2,274 M per HF metadata)
Architecture
Hybrid Gated-DeltaNet (dense FFN) — 24 LM layers (18 DeltaNet + 6 full-attention), ViT vision encoder
Type
Pre-trained base model (not instruction-tuned)
Context
262 144 tokens
Modalities
Text + Vision (early-fusion multimodal)
Key Contributions
Only multimodal… See the full description on the dataset page: https://huggingface.co/datasets/F555/qwen3.5-2b-base-blind-spots.binance-btcusdt-spot-tradesopenwateratlas-spots-climate-species-routes
OpenWaterAtlas — Spots × Climate × Species × Routes
Version 1.0.0 · 2026-06-12 · CC-BY 4.0 · 1,005,290 rows across 8 configs
An integrated open dataset joining 2,928 dive, kitesurf, surf and freedive spots to:
daily climate aggregates (wind, air temperature, precipitation) at each spot,
marine species occurrences (OBIS + GBIF, terrestrial-noise filtered) at each spot,
the global direct-route airline graph (OpenFlights, 34,876 non-stop pairs).
This is the structured-data layer… See the full description on the dataset page: https://huggingface.co/datasets/Anitrovic/openwateratlas-spots-climate-species-routes.spotify-databtcusdt_spot_1m_05_2021_to_03_2026
BTCUSDT Spot 1-Minute OHLCV (May 2021 - Mar 2026)
Overview
1-minute OHLCV candlestick data for the BTC/USDT spot pair on Binance, covering May 1, 2021 to February 28, 2026.
Rows: 2,541,600
Completeness: 100.00%
Sources
Period
Source
Notes
Full dataset
Binance Data Collection
Monthly kline ZIPs
2021-08-13 02:00-06:29
Bybit API
270 bars filled from Bybit BTCUSDT spot (Binance maintenance)
2021-09-29 07:00-08:59
Bybit API
120 bars… See the full description on the dataset page: https://huggingface.co/datasets/addyAIMLprojects/btcusdt_spot_1m_05_2021_to_03_2026.xrpusdt_spot_1m_05_2021_to_03_2026
XRPUSDT Spot 1-Minute OHLCV Dataset
1-minute OHLCV candlestick data for the XRP/USDT spot pair on Binance,
covering May 1, 2021 to February 28, 2026.
Rows: 2,541,600
Completeness: 100.00%
Time Range: May 1, 2021 — February 28, 2026
Columns
Column
Type
Description
timestamp
datetime64[ns]
Candle open time (UTC)
open
float64
Opening price (USDT)
high
float64
Highest price in the candle
low
float64
Lowest price in the candle
close
float64
Closing… See the full description on the dataset page: https://huggingface.co/datasets/Mindbyte-89/xrpusdt_spot_1m_05_2021_to_03_2026.Selena-Gomez-With-Lyrics-And-Spotify-Audio-Featuresbtcusdt_spot_1m_03_2023_to_03_2026
BTCUSDT Spot 1-Minute OHLCV (Mar 2023 - Mar 2026)
Overview
1-minute OHLCV candlestick data for the BTC/USDT spot pair on Binance, covering March 1, 2023 to February 28, 2026.
Rows: 1,578,240
Completeness: 100.00%
Sources
Period
Source
Notes
Full dataset
Binance Data Collection
Monthly kline ZIPs
2023-03-24 12:30-13:59
Bybit API
90 bars filled from Bybit BTCUSDT spot (Binance was in maintenance)
Columns
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/Torch-Trade/btcusdt_spot_1m_03_2023_to_03_2026.pnp_dual_camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "tm",
"total_episodes": 30,
"total_frames": 17862,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/moom-spotter/pnp_dual_cam.spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!
Column… See the full description on the dataset page: https://huggingface.co/datasets/sfiore/spotify-tracks-dataset.
