datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.free-music-archive-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.arc-voicesamples-generatedarc-speeches-refinedfree-music-archive-medium
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-medium.free-music-archive-small
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-small.free-music-archive-large
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-large.ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.free-music-archive-commercial-16khz-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-commercial-16khz-full.free-music-archive-retrieval
FMAR: A Dataset for Robust Song Identification
Authors: Ryan Lee, Yi-Chieh Chiu, Abhir Karande, Ayush Goyal, Harrison Pearl, Matthew Hong, Spencer Cobb
Overview
To improve copyright infringement detection, we introduce Free-Music-Archive-Retrieval (FMAR), a structured dataset designed to test a model's capability to identify songs based on 5-second clips, or queries. We create adversarial queries to replicate common strategies to evade copyright infringement detectors… See the full description on the dataset page: https://huggingface.co/datasets/ml-ryanlee/free-music-archive-retrieval.archi_rutul_asr
Data Sources
Archi
@misc{kibrik2007Archi,
title = {Archi text corpus (1.0)},
author = {Kibrik, Aleksandr E. and Kodzasov, Sandro V. and Olovyannikova, Irina P. and Samedov, Dzhalil S. and Daniel, Michael and Khoroshkina, Anna and Arkhipov, Alexandre},
year = {2007},
url = { https://doi.org/10.5281/zenodo.8247597}
}
Kina Rutul
@misc{alekseevaetal2024,
title = {Dictionary of Kina Rutul},
author = {Alekseeva, Anastasia and Beklemishev, Nikita and Daniel… See the full description on the dataset page: https://huggingface.co/datasets/mahesh27/archi_rutul_asr.cmu-arctic
CMU Arctic Dataset
speech_accent_archive_synthnhk-archive-audio-30s
NHK Archives Audio 30s
This is a Japanese speech corpus derived from NHK Archives Audio. Audio from public NHK Archives records was segmented into clips of up to 30 seconds using voice activity detection.
The dataset contains 137,594 accepted clips, totaling 1,068.96 hours. Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text contains LLM-assisted corrections based on the transcript and available source title and description.
This… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio-30s.french_tv_media_dataset_2026
Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus
Résumé (Abstract)
Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/Archime/french_tv_media_dataset_2026.speech_accent_archive_englishnhk-archive-audio
NHK Archives Audio
NHK Archives Audio is a Japanese audio corpus built from records in the public NHK Archives search service. It contains the audio tracks of archive video and audio records together with titles, descriptions, genres, broadcast metadata, regions, source pages, direct stream URLs, and duration metadata.
The source streams were converted to 16 kHz mono FLAC and embedded directly in Parquet files for use with the Hugging Face Dataset Viewer. This dataset does not… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio.schale-archive-mirror
Schale Archive
Overview
This repository contains Blue Archive game assets.
Known Missing Characters Spine
NPC
GSC President
Yume
Nao
Nagusa
Nyanten-maru
Pei
Reizyo
A.R.O.N.A
How to Properly Import Character to Spine Editor
First, you need Spine Editor at least v3.8.x and later.
Create a new project.
Delete the default skeleton in Hierarchy on the right.
Click Spine logo on top left, then select Import Data.
Select… See the full description on the dataset page: https://huggingface.co/datasets/Hutao0514/schale-archive-mirror.speech_accent_archive_otheraijockey-public-corpusDatasetKirAll250-hf
DatasetKirAll250-hf
CSV структура:
audio_file|text|speaker_name
wavs/имя_файла.wav|...|...
ruby
Copy code
Файлы шукаюцца ў /content/DatasetKirAll250/dataset-1/wavs; калі ў CSV пазначана .wav, але існуе .mp3 (ці .flac),
скрыпт аўтаматычна падмяняе пашырэнне і знаходзіць файл.
Палі:
audio — аўдыяфайл (аўтаматычна загружаецца як asset у Hub).
text — транскрыпцыя.
speaker — імя/ID спікера.
Створана: 2025-09-27.
sinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.l2-arctic-manual-v5.0-16k
l2-arctic-manual-v5.0-16k
This dataset is a prepared derivative of L2-ARCTIC v5.0 that keeps only
the manually annotated material and converts the audio to 16 kHz mono FLAC.
It is designed to plug into the current peacock-asr training code, which
can consume a Hugging Face dataset with audio plus phonemes.
Included splits
train: 1800 rows, 1.84 hours
validation: 899 rows, 0.94 hours
test: 900 rows, 0.88 hours
suitcase: 22 rows, 0.44 hours
The scripted subset uses the… See the full description on the dataset page: https://huggingface.co/datasets/chikingsley/l2-arctic-manual-v5.0-16k.hindi-whisper-chunks
Hindi Whisper Chunks
Preprocessed, feature-extracted audio chunks and labels used to fine-tune ArchCoder/whisper-small-hindi-lora, a LoRA adaptation of Whisper-small for Hindi speech recognition.
Dataset Summary
Raw Hindi audio recordings (approximately 12 minutes each) were segmented into short, Whisper-compatible chunks and converted into model-ready features. This dataset is the output of that preprocessing pipeline: Whisper-format log-mel filterbank features… See the full description on the dataset page: https://huggingface.co/datasets/ArchCoder/hindi-whisper-chunks.aliaksandr-serzhputouski-kazki-i-apaviadanni-belarusau-slutskaga-pavetu-iury-zhy
Казкі і апавяданні беларусаў Слуцкага павету
Metadata
Author: Аляксандр Сержпутоўскі
Title: Казкі і апавяданні беларусаў Слуцкага павету
Narrator: Юры Жыгамонт
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/aliaksandr-serzhputouski-kazki-i-apaviadanni-belarusau-slutskaga-pavetu-iury-zhy.chiptune_musicthis dataset was extracted from youtube with videos under cc (creatives common) tag so theyre are free use
music-fingerprint-dataset
Neural Audio Fingerprint Dataset
(c) 2021 by Sungkyun Chang
https://github.com/mimbres/neural-audio-fp
This dataset includes all music sources, background noise and impulse-reponses
(IR) samples that have been used in the work ["Neural Audio Fingerprint for
High-specific Audio Retrieval based on Contrastive Learning"]
(https://arxiv.org/abs/2010.11910).
Format:
16-bit PCM Mono WAV, Sampling rate 8000 Hz
Description:
/
fingerprint_dataset_icassp2021/… See the full description on the dataset page: https://huggingface.co/datasets/arch-raven/music-fingerprint-dataset.free-music-archive-medium
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/Trkaga/free-music-archive-medium.product-bench
Product Bench
31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a laptop comparison shopping assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a laptop comparison shopping assistant helping a customer evaluate, compare, and order laptops. The conversation features multi-intent turns… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/product-bench.ARCADE-full
ARCADE: Arabic Radio Corpus for Audio Dialect Evaluation
ARCADE is a city-scale corpus of Arabic radio speech designed for fine-grained dialect identification. The dataset contains 6,907 annotations for 3,790 unique audio segments collected from radio streams spanning 58 cities across 19 Arab countries.
Dataset Description
Each 30-second audio clip is annotated with:
City and Country: Fine-grained geographic labels at the city level
MSA or Dialect: Whether the speech is… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/ARCADE-full.
