datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
podcasts_spotifyspotify-data-960kSpotify_Audio_features_2.3Membeat_45m_spotify_tracks
Embeat 45M Spotify Tracks
A large-scale music metadata dataset containing 45 million Spotify tracks, combined with the Spotify metadata from Anna's Archive and artist genres from Every Noise at Once.
GitHub project: https://github.com/gdstudio-org/Embeat
Dataset Preview
>>> from datasets import load_from_disk
>>> ds = load_from_disk("GD-Studio/embeat_45m_spotify_tracks")
>>> ds
Dataset({
features: ['track_id', 'track_name', 'isrc', 'popularity', 'explicit'… See the full description on the dataset page: https://huggingface.co/datasets/GD-Studio/embeat_45m_spotify_tracks.spotify_audio_features
Spotify Tracks & Audio Features Dataset
Overview
This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research.
Data Source
The raw data for this dataset was originally gathered and hosted by Anna's Archive.… See the full description on the dataset page: https://huggingface.co/datasets/0xaf/spotify_audio_features.spotify-genresSpotify genres scraped from https://everynoise.com/everynoise1d.cgi?scope=all
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
dataset_info:
features:
- name: genre_name
dtype: string
- name: genre_slug
dtype: string
- name: playlist_url
dtype: string
- name: description
dtype: string
splits:
- name: train
num_bytes: 1047789
num_examples: 6276
download_size: 577290
dataset_size: 1047789
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/devxpy/spotify-genres.Spotify-Africa-Dataset
Spotify-Africa Music Dataset 🎵🌍
A comprehensive, research-grade dataset documenting African music from Spotify spanning 1,600+ tracks, 650+ artists, and 67 years of musical history (1958-2025).
Dataset Summary
This dataset provides rich metadata about African music across multiple genres, regions, and time periods. It includes track-level information, artist metadata, temporal trends, regional summaries, and network relationships. The data was collected via the Spotify… See the full description on the dataset page: https://huggingface.co/datasets/Kossisoroyce/Spotify-Africa-Dataset.spotify-track-featureshubert_process_filter_spotifyspotifyspotify-lyrics-valence-originalSpotify-Africa-Dataset
Spotify-Africa Music Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Spotify-Africa-Dataset.fqw-spotify-lyricsspotify-trainingplaylists/ (6.6M rows, 2 files)
┌─────────────────────┬────────┬─────────────────────────────┐
│ Column │ Type │ Sample │
├─────────────────────┼────────┼─────────────────────────────┤
│ rowid │ int64 │ 1 │
│ id │ string │ "37i9dQZF1DXcBWIGoYBM5M" │
│ snapshot_id │ string │ "MTczODcwNTA4MCww..." │
│ fetched_at │ int64 │ 1703456789000 │
│ name │… See the full description on the dataset page: https://huggingface.co/datasets/Cossale/spotify-training.spotify_musichubert_processed_spotifywav2vec2_processed_spotify-nonstratifiedspotify-valence-50k-trainset1spotify
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AnaLFDias/spotify.Spotify_Million_Songhttps://www.kaggle.com/
spotify_dataset_whisperspotify-valence-trainset3spotify_raw_small_2spotify-million-song-dataset-descriptionsraw_train_dev_test_spotifyspotify-valence-50k-trainset2Spotify10Testwav2vec_filter_wo_processing_spotifyspotify_raw_small_1spotify-rekomendasi-data
