datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!
Column… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/spotify-tracks-dataset.spooky-author-identificationspotify-tracksnangang_sports_centerspotify-tracks-lite
Context
This dataset consists of 24000 tracks from 30 genres, and is a shrunk version of maharshipandya/spotify-tracks-dataset dataset. All non-heuristic data is cut and cleaned for better usability and performance.
All data taken from Spotify API and is open source.
This dataset can be used to train prediction models based on user preferences, or categorise tracks by corresponding heuristic.
Column Description
danceability: Danceability describes how suitable a track is… See the full description on the dataset page: https://huggingface.co/datasets/engels/spotify-tracks-lite.sportsbookish-daily-odds
SportsBookISH Daily Kalshi vs Sportsbook Odds
Real-time pricing snapshot comparing Kalshi event-contract probabilities against US sportsbook consensus across nine sports.
Description
Daily-refreshed JSON / CSV export of every active Kalshi market alongside the de-vigged book median across 13+ US sportsbooks. Covers golf (PGA Tour), NFL, NBA, MLB, NHL, EPL, MLS, UEFA Champions League, and FIFA World Cup.
Source
Live data plane:
JSON:… See the full description on the dataset page: https://huggingface.co/datasets/kennyhyder/sportsbookish-daily-odds.gnss-jamming-spoofing-detection
GNSS Jamming & Spoofing Detection Dataset
A physics-informed synthetic dataset for detecting GPS/GNSS cyber-attacks
(Jamming and Spoofing) from satellite-signal features. Built for the
GNSS Guardian project — Introduction to Data Science final project.
Overview
14,850 samples across 450 scenarios × 33 time-steps each
3 balanced classes: Normal / Jamming / Spoofing (4,950 each)
26 columns: multi-constellation signal features + attack metadata + text descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Omrilevi123/gnss-jamming-spoofing-detection.top-hits-spotifyFISH_spots
FISH_spots Dataset
The manually verified in situ hybridization fluorescence images and point coordinate dataset.
This dataset contains images and annotations for the task of single-molecule fluorescence in situ hybridization (FISH) spot detection, supporting 2D, 3D, and simulated noisy data. The structure is designed for deep learning model development, training, and evaluation.
Directory Structure
FISH_spots/
├── 2d/
│ ├── csv/
│ ├── image/
│ ├── image_raw/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/GangCaoLab/FISH_spots.spotify-million-song-dataset
Dataset Card for Spotify Million Song Dataset
Dataset Summary
This is Spotify Million Song Dataset. This dataset contains song names, artists names, link to the song and lyrics. This dataset can be used for recommending songs, classifying or clustering songs.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data… See the full description on the dataset page: https://huggingface.co/datasets/vishnupriyavr/spotify-million-song-dataset.SpotifyDataavs-spot
Dataset Card for AVS-Spot Benchmark
This dataset is associated with the paper: "Understanding Co-Speech Gestures in-the-wild"
📝 ArXiv: https://arxiv.org/abs/2503.22668
🌐 Project page: https://www.robots.ox.ac.uk/~vgg/research/jegal
💻 Code: https://github.com/Sindhu-Hegde/jegal
We present JEGAL, a Joint Embedding space for Gestures, Audio and Language. Our semantic gesture representations can be used to perform multiple downstream tasks such as cross-modal retrieval… See the full description on the dataset page: https://huggingface.co/datasets/sindhuhegde/avs-spot.Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.spotify-lyrics
Dataset Card for Spotify Million Song Dataset
Dataset Summary
This is Spotify Million Song Dataset. This dataset contains song names, artists names, link to the song and lyrics. This dataset can be used for recommending songs, classifying or clustering songs.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data… See the full description on the dataset page: https://huggingface.co/datasets/vancenceho/spotify-lyrics.spotify-tracks-datasetSpotifyLyrics001spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
spotify_popular_tracksspotifymodelspotify-artistsspotify-dataspotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!
Column… See the full description on the dataset page: https://huggingface.co/datasets/sfiore/spotify-tracks-dataset.spotify-lyrics-clean
Spotify Million Song Lyrics Cleaned
CSV of deduplicated, normalized lyrics aligned to the Million Song–style raw lyrics dump used in the viral-content-predictor project. Each row is one (artist, song) identity after normalization; lyrics are cleaned text suitable for TF‑IDF, retrieval, or joining to Spotify metadata via artist_norm / title_norm.
File
File
Role
lyrics_cleaned.csv
One row per normalized (artist_norm, title_norm); includes raw display columns and… See the full description on the dataset page: https://huggingface.co/datasets/vancenceho/spotify-lyrics-clean.spotify_datasetspotify-million-song
Dataset Card for Dataset Name
A dataset containing songs, artists names, link to song and lyrics
Dataset Details
Dataset retrieved from https://www.kaggle.com/datasets/notshrirang/spotify-million-song-dataset
Dataset Description
This is Spotify Million Song Dataset. This dataset contains song names, artists names, link to the song and lyrics. This dataset can be used for recommending songs, classifying or clustering songs.
Curated by: SHRIRANG MAHAJAN… See the full description on the dataset page: https://huggingface.co/datasets/sebastiandizon/spotify-million-song.SportsMetrics
SportsMetrics
Benchmark data to evaluate numerical reasoning and information fusion of LLMs.
SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL'24), Bangkok, Thailand. Arxiv Paper
Usage
from datasets import load_dataset
def get_task(domain… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/SportsMetrics.JPOP-Sampled-Spotifysports-basketball-coach
This Dialogue
Comprised of fictitious examples of dialogues between a basketball coach and the players on the court during a game. Check out the example below:
"id": 1,
"description": "Motivating the team",
"dialogue": "Coach: Let's give it our all, team! We've trained hard for this game, and I know we can come out on top if we work together."
How to Load Dialogues
Loading dialogues can be accomplished using the fun dialogues library or Hugging Face datasets library.… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/sports-basketball-coach.Spotify_Songs_with_SoundCloud_linksspotify-hit-prediction-analysis
Your browser does not support the video tag.
🎵 Spotify Hit Prediction - Exploratory Data Analysis (EDA)
Project Overview
This project analyzes audio features from Spotify to predict track popularity. Using a sample of 2,000 tracks, I explored how technical attributes like energy and danceability relate to a song's success.
🔍 Research Questions & Insights
I addressed several key questions during the EDA:
Is the data balanced? I analyzed the ratio of… See the full description on the dataset page: https://huggingface.co/datasets/Ohad777/spotify-hit-prediction-analysis.
