datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stack_cubesjam-alt-lines
Jam-ALT Lines
Jam-ALT Lines is a line-level version of the Jam-ALT lyrics transcription dataset.
Unlike Jam-ALT, this dataset contains one audio segment for each lyrics line, facilitating research that considers each line as a separate unit.
[!tip]
See the Jam-ALT project website for details and the JamendoLyrics community for related datasets.
Dataset flavors
Lyrics lines may overlap in time, which makes it impossible to have a one-to-one correspondence between… See the full description on the dataset page: https://huggingface.co/datasets/cublya/jam-alt-lines.TurkishVoiceDataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/cubukcum/TurkishVoiceDataset.jam-alt
Jam-ALT: A Readability-Aware Lyrics Transcription Benchmark
Jam-ALT is a revision of the JamendoLyrics dataset (79 songs in 4 languages), intended for use as an automatic lyrics transcription (ALT) benchmark.
It has been published in the ISMIR 2024 paper (full citation below): 📄 Lyrics Transcription for Humans: A Readability-Aware Benchmark 👥 O. Cífka, H. Schreiber, L. Miner, F.-R. Stöter 🏢 AudioShake
The lyrics have been revised according to the newly compiled annotation… See the full description on the dataset page: https://huggingface.co/datasets/cublya/jam-alt.audio_swedish_2_dataset_cleanedjamendolyrics
JamendoLyrics MultiLang dataset for lyrics research
A dataset containing 79 songs with different genres and languages along with lyrics that
are time-aligned on a word-by-word level (with start and end times) to the music.
[!note]
Note: The dataset is primarily intended as an automatic lyrics alignment (ALA) benchmark.
For lyrics transcription, please see the Jam-ALT
dataset, which contains a revised version of the lyrics, better suited as a reference for the transcription task.… See the full description on the dataset page: https://huggingface.co/datasets/cublya/jamendolyrics.audio_swedish_2_datasetkanata1
