datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spotify-songsspotify_songsReadme
Dataset Description:
This dataset is brought from kaggle: "30000 Spotify Songs". The dataset contains both numeric and categorical variables describing songs available on Spotify. It includes musical characteristics such as danceability, energy, loudness, valence, tempo, and duration, as well as metadata like artist, album, and genre.
Research Question:
What song characteristics make a track more popular on Spotify?
Target Variable:
The target variable is track_popularity, which… See the full description on the dataset page: https://huggingface.co/datasets/uleeberber/spotify_songs.spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
Songsidol-songs-jpEnglish version follows the Japanese text.
IdolSongsJp: アイドルグループ楽曲スタイルにもとづく音楽コーパス
日本のアイドルグループを模した 15 のオリジナル楽曲からなるコーパスです。
ティザームービー
コーパス構成
楽曲と歌唱者
15 楽曲のうち、8 楽曲が女性グループ、7 曲が男性グループによる楽曲です。
女性歌唱者は 10 名、男性歌唱者は 8 名であり、歌唱メンバーは楽曲によって異なります。
作曲者・編曲者・作詞者はすべてアイドルグループに楽曲提供経験を持つプロフェッショナルです。各歌唱者はセミプロフェッショナルもしくはプロフェッショナルのボーカリストです(実際のアイドルではありません)。
楽曲の一覧と制作者を下記に示します。楽曲長は切り捨てです。BPM は代表値を示しています。
楽曲 ID
楽曲名
作詞
作曲
編曲
楽曲長
BPM
歌唱者数
f01-intro_juice
いんとろじゅーす
ハイジナカムラ… See the full description on the dataset page: https://huggingface.co/datasets/imprt/idol-songs-jp.songsspotify-songsACEStep-Songssongs generated by ACE-Step
score_lyrics field represnets the scores given by gpt-4o for the lyrics(1-10, only those >=8 are preserved), -1 for instrumental piece
the full lyrics-tags-score dataset is at https://huggingface.co/datasets/Yi3852/lyrics-tags_gen
more info see https://github.com/ace-step/ACE-Step/issues/313
Citation
@misc{jiang2025advancingfoundationmodelmusic,
title={Advancing the Foundation Model for Music Understanding},
author={Yi Jiang and Wei Wang… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/ACEStep-Songs.formatted_songsclean-songs-lyrics-dataset
Clean Songs Lyrics Dataset
1.53M+ clean songs lyrics with songs titles and artists names
Dataset info
This is the combined, deduped, cleaned, and sanitized aggregation of three large lyrics datasets
Each lyric was deduplicated
Each lyric was checked to be in range of 256 bytes <-> 8192 bytes
Each lyric was checked for profanities with alt-profanity-check
Each lyric was ASCII sanitized for conistency… See the full description on the dataset page: https://huggingface.co/datasets/asigalov61/clean-songs-lyrics-dataset.pop909-with-songskaraoke_songs_long
Karaoke Songs Long Dataset (karaoke_songs_long)
A comprehensive collection of instrumental karaoke versions of popular songs in WAV format, spanning multiple genres and artists. Includes a captions.csv file with metadata for each track.
Dataset Description
This dataset contains hundreds of instrumental karaoke tracks (vocals removed/downmixed) in high-quality WAV format. The songs cover a wide range of artists — from Adele and Taylor Swift to Queen, Elvis Presley… See the full description on the dataset page: https://huggingface.co/datasets/edwixx/karaoke_songs_long.original-songs
Dataset Card for "original-songs" (Audio + análisis DSP)
Dataset Summary
Dataset pequeño de canciones originales creadas con IA, cada una con su WAV,
letra transcrita automáticamente (Whisper) y un análisis DSP completo (tempo,
tonalidad, loudness, features perceptuales) además de detección de contenido
explícito. Pensado para quien quiera mejorar modelos open source: extracción
de features musicales, clasificación de audio, transcripción y moderación de
letras.… See the full description on the dataset page: https://huggingface.co/datasets/arnauquest/original-songs.English_French_Songs_Lyrics_Translation_Original
Original Songs Lyrics with French Translation
Dataset Summary
Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French.
Details of the number of songs by language of origin can be found in the table below:
Original language
Number of songs
en
75786
fr
18486
es
1743
it
803
de
691
sw
529
ko
193
id
169
pt
142
no
122
fi
113
sv
70
hr
53
so
43
ca
41
tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.uta-net-songs
Dataset Details
Dataset Description
This dataset contains a processed version of a web scrape I did for uta-net. The raw data is available for download at here.
Uta-Net site mainly lists songs that have been released in Japan officially (Anime OP/EDs) up to 2023-03.
Curated by: KaraKaraWitch
Shared by: KaraKaraWitch
Language(s) (NLP): JA
License: Not Disclosed / Unsure
Stuff not in this dataset:
Character Songs for Anime
Doujin/Indie Works
Dataset Sample… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/uta-net-songs.ai-generated-songshit-songs-private-reserve
Hit Songs [Private Reserve]
Select hand-picked greatest music hit songs
License
Strictly CC BY-NC-SA !!!
Music sources
YouTube
SoundCloud
Project Los Angeles
Tegridy Code 2026
Spotify_Songs_with_SoundCloud_linksfrench_rap_songssongssongssong_structure_with_testabdulszz_spotify-most-streamed-songs
Spotify Most Streamed Songs
Unveiling Streaming: A Comprehensive Analysis of Spotify’s Most Streamed Songs
Dataset Info
Source: Kaggle
Original Size: 0.06 MB
Kaggle Downloads: 25,259
Files: 1
Files
Spotify Most Streamed Songs.csv
Mirrored from Kaggle
ai-generated-songs2Songs-for-transcriptionBollywood_songs
Dataset Card for "Bollywood_songs"
More Information needed
5M-Songs-Lyrics
Dataset Summary
This dataset contains 50 million rows of song lyrics sourced from a public Kaggle dataset. It has been preprocessed into an instruction–label format suitable for training or fine-tuning generative language models, particularly for music lyric generation tasks.
Each row is designed to guide a model to generate song verses in the style of a specific artist and genre, with corresponding real lyric snippets as ground truth.
Supported Tasks and Benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/rajtripathi/5M-Songs-Lyrics.LYRICAL_Mix_SilverAgePoets_Songs_RuVerses_SFT
SilverAgePoets.com & RuVERSES.com Russian-English Bilingual Poetry Library
A dataset of Eastern European and Soviet poetry and song lyrocs from https://RuVerses.com/, with Russian-language sources and English translations.
This variant of the dataset combines a revised and somewhat pre-filtered version of the RuVerses collection dataset + the entirety of the SFT version of our LYRICAL dataset.
Featuring a present (c. late 2025) state of the RuVerses archive, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/LYRICAL_Mix_SilverAgePoets_Songs_RuVerses_SFT.CEBThe dataset for bias evaluation of LLMs. Github: https://github.com/SongW-SW/CEB
Lyrical_rus2eng_ORPOv5.1_SongsPoems_MeteredTranslations_csv
Meaning+Meter-Matched Russian & Soviet Poems + Songs
Manually Translated by a Poet-Translator from Russian to English
Translations herein faithfully adapt the Source Lyrics' Metered/Rhythmic/Rhyming Patterns
NEWLY EDITED VARIANT 5.1: 1775 rows/items
Re-balanced, refined, standardized, and substantially expanded.
CSV version
Manually translated to English by Aleksey Calvin, with a painstaking effort to cross-linguistically reproduce source texts'… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_rus2eng_ORPOv5.1_SongsPoems_MeteredTranslations_csv.
