datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!
Column… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/spotify-tracks-dataset.FISH_spots
FISH_spots Dataset
The manually verified in situ hybridization fluorescence images and point coordinate dataset.
This dataset contains images and annotations for the task of single-molecule fluorescence in situ hybridization (FISH) spot detection, supporting 2D, 3D, and simulated noisy data. The structure is designed for deep learning model development, training, and evaluation.
Directory Structure
FISH_spots/
├── 2d/
│ ├── csv/
│ ├── image/
│ ├── image_raw/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/GangCaoLab/FISH_spots.btcusdt_spot_1m_03_2023_to_12_2025spotify-songsbinance-top50-spot-v1
Binance Top 50 Backtesting Dataset
Built at: 2026-05-21T11:53:45.468482+00:00
Parameters
Lookback: 1 days
Top N: 3
Trade Types: spot, um
Data Types: klines, aggTrades
Build Status
SPOT: 3 symbols
klines: 3/3 healthy
aggTrades: 3/3 healthy
UM: 3 symbols
klines: 3/3 healthy
aggTrades: 3/3 healthy
fundingRate: 3/3 healthy
Spotlight-VideoGen-Errors
Spotlight Dataset
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
Aditya Chinchure, Sahithya Ravi, Pushkar Shukla, Vered Shwartz, Leonid Sigal
🎉 Accepted to ECCV 2026
🌐 Project Page
Summary
Spotlight is a benchmark for evaluating whether Vision Language Models (VLMs) can precisely
localize and explain errors in AI-generated videos. It contains 600 videos generated by
three state-of-the-art Text-to-Video (T2V) models —… See the full description on the dataset page: https://huggingface.co/datasets/UBC-ViL/Spotlight-VideoGen-Errors.messy_pick_object_place_plat_spotPick [object from the green box/ egg from the large round plate] and place it in the frying pan.
spotify-tracksvlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
spot-terrain-dataset
Spot Dataset
SpotSFT-200k
SpotSFT-200k: Visual QA Dataset for Geo-localization Alignment
Project Page
Dataset Description
SpotSFT-200k is a large-scale multimodal instruction-tuning dataset comprising approximately 200,000 image-text pairs. It is designed for the Supervised Fine-Tuning (SFT) stage of the SpotAgent framework (Stage 1).
Unlike the subsequent SpotAgenticCoT dataset which focuses on complex tool use and reasoning, SpotSFT-200k aims to:
Inject Basic World Knowledge: Align the… See the full description on the dataset page: https://huggingface.co/datasets/jiafr1802/SpotSFT-200k.spotify_audio_features
Spotify Tracks & Audio Features Dataset
Overview
This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research.
Data Source
The raw data for this dataset was originally gathered and hosted by Anna's Archive.
Original Blog Post:… See the full description on the dataset page: https://huggingface.co/datasets/ozefe/spotify_audio_features.spotify-tracks-lite
Context
This dataset consists of 24000 tracks from 30 genres, and is a shrunk version of maharshipandya/spotify-tracks-dataset dataset. All non-heuristic data is cut and cleaned for better usability and performance.
All data taken from Spotify API and is open source.
This dataset can be used to train prediction models based on user preferences, or categorise tracks by corresponding heuristic.
Column Description
danceability: Danceability describes how suitable a track is… See the full description on the dataset page: https://huggingface.co/datasets/engels/spotify-tracks-lite.top-hits-spotifyhab_spot_arm
Boston Dynamics Spot + Arm
Simulation model (URDF) of Boston Dynamics Spot robot + manipulator arm module for use in habitat-sim.
License Information
See LICENSE.txt for more details.
Original "urdf/hab_spot_arm.urdf" and all assets referenced there-in are provided courtesy of Boston Dynamics, all rights reserved.
All other assets represent derivative work of said authors.
Written permission has been acquired for redistribution of these assets with attribution.
spotify_songsReadme
Dataset Description:
This dataset is brought from kaggle: "30000 Spotify Songs". The dataset contains both numeric and categorical variables describing songs available on Spotify. It includes musical characteristics such as danceability, energy, loudness, valence, tempo, and duration, as well as metadata like artist, album, and genre.
Research Question:
What song characteristics make a track more popular on Spotify?
Target Variable:
The target variable is track_popularity, which… See the full description on the dataset page: https://huggingface.co/datasets/uleeberber/spotify_songs.spotify-huge-track-analysis-dataset
Spotify Track Analysis Dataset
General Description
This dataset provides a large-scale, research-oriented analytical representation of Spotify music data.
It is centered on tracks as musical recordings (track_id), while preserving explicit artist attribution as defined by Spotify’s native credit model.
Each row corresponds to a track–artist association, identified by:
a Spotify track identifier (track_id)
a credited artist name (artist_name)
A single track may appear on… See the full description on the dataset page: https://huggingface.co/datasets/GildasLeDrogoff/spotify-huge-track-analysis-dataset.spotify_audio_features_partitionedb-FLAIR-spot
b-FLAIR-spot: bi-temporal extension of FLAIR in SPOT-6/7 modality
Dataset Description
b-FLAIR-spot is a temporal extension of the FLAIR dataset [1], mirroring b-FLAIR in SPOT-6/7 modality, focused on land cover classification in France. The dataset provides bi-temporal satellite image pairs with single-temporal semantic annotations.
Project page: https://xavibou.github.io/CDviaWTS/
Dataset Summary
Task: Semantic change detection via weak temporal… See the full description on the dataset page: https://huggingface.co/datasets/elliotvincent/b-FLAIR-spot.ethusdt_spot_1m_05_2021_to_03_2026
ETHUSDT Spot 1-Minute OHLCV (May 2021 - Mar 2026)
Overview
1-minute OHLCV candlestick data for the ETH/USDT spot pair on Binance, covering May 1, 2021 to February 28, 2026.
Rows: 2,541,600
Completeness: 100.00%
Sources
Period
Source
Notes
Full dataset
Binance Data Collection
Monthly kline ZIPs
2021-08-13 02:00-06:29
Bybit API
270 bars filled from Bybit ETHUSDT spot (Binance maintenance)
2021-09-29 07:00-08:59
Bybit API
120 bars filled from… See the full description on the dataset page: https://huggingface.co/datasets/Torch-Trade/ethusdt_spot_1m_05_2021_to_03_2026.spotify-million-song-dataset
Dataset Card for Spotify Million Song Dataset
Dataset Summary
This is Spotify Million Song Dataset. This dataset contains song names, artists names, link to the song and lyrics. This dataset can be used for recommending songs, classifying or clustering songs.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data… See the full description on the dataset page: https://huggingface.co/datasets/vishnupriyavr/spotify-million-song-dataset.SpotifyLyrics001SpotSound-Bench
SpotSound-Bench: A 'Needle-in-a-Haystack' Evaluation for Audio Temporal Grounding
Benchmark Summary
SpotSound-Bench is a challenging temporal grounding benchmark designed to evaluate Large Audio-Language Models (ALMs).
Existing benchmarks for audio temporal grounding often feature high ratios of target-window duration to full audio clip duration, which fail to simulate real-world scenarios where short events are obscured by dense background sounds. To bridge… See the full description on the dataset page: https://huggingface.co/datasets/Loie/SpotSound-Bench.spotify-lyrics
Dataset Card for Spotify Million Song Dataset
Dataset Summary
This is Spotify Million Song Dataset. This dataset contains song names, artists names, link to the song and lyrics. This dataset can be used for recommending songs, classifying or clustering songs.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data… See the full description on the dataset page: https://huggingface.co/datasets/vancenceho/spotify-lyrics.avs-spot
Dataset Card for AVS-Spot Benchmark
This dataset is associated with the paper: "Understanding Co-Speech Gestures in-the-wild"
📝 ArXiv: https://arxiv.org/abs/2503.22668
🌐 Project page: https://www.robots.ox.ac.uk/~vgg/research/jegal
💻 Code: https://github.com/Sindhu-Hegde/jegal
We present JEGAL, a Joint Embedding space for Gestures, Audio and Language. Our semantic gesture representations can be used to perform multiple downstream tasks such as cross-modal retrieval… See the full description on the dataset page: https://huggingface.co/datasets/sindhuhegde/avs-spot.dumpster-diving-spots
Dataset Card for "Dataset of Dumpster Diving Spots"
Community-collected dumpster diving spots and their ratings from https://www.dumpstermap.org (or better https://dumpstermap.herokuapp.com/dumpsters).
Updated monthly using https://github.com/Hitchwiki/dumpster-archiver.
Dataset Details
Dataset Sources
Repository: https://github.com/Debakel/Dumpstermap
Demo: https://www.dumpstermap.org/
Uses
gift and sharing econonmy… See the full description on the dataset page: https://huggingface.co/datasets/Hitchwiki/dumpster-diving-spots.Spotify_Music_Analytics_and_Popularity_Predictionpodcasts_spotifySpotifyDataspot-the-diff
