datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.free-music-archive-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.free-music-archive-medium
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-medium.free-music-archive-small
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-small.free-music-archive-large
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-large.ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.free-music-archive-commercial-16khz-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-commercial-16khz-full.free-music-archive-retrieval
FMAR: A Dataset for Robust Song Identification
Authors: Ryan Lee, Yi-Chieh Chiu, Abhir Karande, Ayush Goyal, Harrison Pearl, Matthew Hong, Spencer Cobb
Overview
To improve copyright infringement detection, we introduce Free-Music-Archive-Retrieval (FMAR), a structured dataset designed to test a model's capability to identify songs based on 5-second clips, or queries. We create adversarial queries to replicate common strategies to evade copyright infringement detectors… See the full description on the dataset page: https://huggingface.co/datasets/ml-ryanlee/free-music-archive-retrieval.speech_accent_archive_synthnhk-archive-audio-30s
NHK Archives Audio 30s
This is a Japanese speech corpus derived from NHK Archives Audio. Audio from public NHK Archives records was segmented into clips of up to 30 seconds using voice activity detection.
The dataset contains 137,594 accepted clips, totaling 1,068.96 hours. Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text contains LLM-assisted corrections based on the transcript and available source title and description.
This… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio-30s.speech_accent_archive_englishspeech_accent_archive_othernhk-archive-audio
NHK Archives Audio
NHK Archives Audio is a Japanese audio corpus built from records in the public NHK Archives search service. It contains the audio tracks of archive video and audio records together with titles, descriptions, genres, broadcast metadata, regions, source pages, direct stream URLs, and duration metadata.
The source streams were converted to 16 kHz mono FLAC and embedded directly in Parquet files for use with the Hugging Face Dataset Viewer. This dataset does not… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio.schale-archive-mirror
Schale Archive
Overview
This repository contains Blue Archive game assets.
Known Missing Characters Spine
NPC
GSC President
Yume
Nao
Nagusa
Nyanten-maru
Pei
Reizyo
A.R.O.N.A
How to Properly Import Character to Spine Editor
First, you need Spine Editor at least v3.8.x and later.
Create a new project.
Delete the default skeleton in Hierarchy on the right.
Click Spine logo on top left, then select Import Data.
Select… See the full description on the dataset page: https://huggingface.co/datasets/Hutao0514/schale-archive-mirror.sinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.Kiko-Archivearxiv_audio_archivedfree-music-archive-medium
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/Trkaga/free-music-archive-medium.free-music-archive-large-filtered2nhk-archive-meta
NHK Archives Metadata
NHK Archives Metadata is a Japanese metadata dataset for audio and video records from the public NHK Archives search service. It provides titles, genres, durations, NHK Archives page URLs, and direct streaming URLs.
The dataset contains metadata and URLs only. It does not contain audio or video files, transcripts, or copied media content.
Purpose
This dataset is intended for research and applications that use Japanese audio and video metadata… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-meta.speech_accent_archive_datasetarchive-bd66a225
Korean multi-speaker speech corpus
46,346 studio-recorded Korean utterances from 8 professional voice performers,
85h 45m in total, each paired with the exact sentence that was read.
Seven of them work from the same five script families, so a large part of the corpus is the
same sentence read by different voices — usable for speech synthesis, voice conversion and
speaker-representation work.
Performers are identified only by opaque codes. Nothing here names a person.
A private… See the full description on the dataset page: https://huggingface.co/datasets/DavidVita/archive-bd66a225.genshin-voice-archivedThis is an English only variant of the simon3000/genshin-voice dataset.
All the credit for this dataset goes to simon3000. Do checkout his work!
gyukaro-research-archivetransatlantic-voice-archive_distille
Distillation brucemacd/transatlantic-voice-archive
Dataset ASR distillé via Cohere Transcribe.
Source : brucemacd/transatlantic-voice-archive
Modèle ASR : cohere-transcribe
Langue ASR : en
Exemples : 1427 (dataset source intégral)
Colonnes :
audio — clip audio (16 kHz)
transcription_base — référence brute du dataset source
transcription_cohere — hypothèse Cohere brute
langue_accent — langue / accent détecté
wer, cer — métriques item (normalisation training_v3, textes stockés… See the full description on the dataset page: https://huggingface.co/datasets/Zeldeo/transatlantic-voice-archive_distille.transatlantic-voice-archive
Transatlantic Voice Archive
Speech clips with aligned transcripts in the transatlantic (Mid-Atlantic)
accent — the clipped, semi-British delivery of 1930s–1960s American newsreel
announcers. Built from public-domain Universal
Newsreels (1929–1967) on the
Internet Archive, intended for finetuning TTS models on the accent.
Dataset statistics
Clips
1427
Total audio
2.04 h (122.2 min)
Average clip
5.14 s
Sample rate
22050 Hz, mono WAV
Source reels… See the full description on the dataset page: https://huggingface.co/datasets/brucemacd/transatlantic-voice-archive.archive_sonores_en_gallo_1
[!CAUTION]
Les données présentes dans ce répertoire ont été recueillies dans le cadre de l'ordonnance n°2021-1518 du 24 novembre 2021 (https://www.legifrance.gouv.fr/jorf/id/JORFTEXT000044362034/#JORFARTI000044362035).C'est-à-dire, que du fait du droit d'auteur, nous les mettons à disposition uniquement pour des chercheurs travaillant dans des instituts de recherche français (INRIA, CNRS, INSERM, etc.) et des universités françaises.
Description
Ce jeu de données fait partie… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/archive_sonores_en_gallo_1.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 1,097
Chunk audio duration: 5.143 hours
Source transcript rows represented: 3,875
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
19
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 2,116
Chunk audio duration: 6.624 hours
Source transcript rows represented: 5,601
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
56
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06.archive_sonores_en_breton_33
[!CAUTION]
Les données présentes dans ce répertoire ont été recueillies dans le cadre de l'ordonnance n°2021-1518 du 24 novembre 2021 (https://www.legifrance.gouv.fr/jorf/id/JORFTEXT000044362034/#JORFARTI000044362035).C'est-à-dire, que du fait du droit d'auteur, nous les mettons à disposition uniquement pour des chercheurs travaillant dans des instituts de recherche français (INRIA, CNRS, INSERM, etc.) et des universités françaises.
Description
Ce jeu de données fait partie… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/archive_sonores_en_breton_33.
