datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.nhk-archive-audio-30s
NHK Archives Audio 30s
This is a Japanese speech corpus derived from NHK Archives Audio. Audio from public NHK Archives records was segmented into clips of up to 30 seconds using voice activity detection.
The dataset contains 137,594 accepted clips, totaling 1,068.96 hours. Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text contains LLM-assisted corrections based on the transcript and available source title and description.
This… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio-30s.sinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.transatlantic-voice-archive_distille
Distillation brucemacd/transatlantic-voice-archive
Dataset ASR distillé via Cohere Transcribe.
Source : brucemacd/transatlantic-voice-archive
Modèle ASR : cohere-transcribe
Langue ASR : en
Exemples : 1427 (dataset source intégral)
Colonnes :
audio — clip audio (16 kHz)
transcription_base — référence brute du dataset source
transcription_cohere — hypothèse Cohere brute
langue_accent — langue / accent détecté
wer, cer — métriques item (normalisation training_v3, textes stockés… See the full description on the dataset page: https://huggingface.co/datasets/Zeldeo/transatlantic-voice-archive_distille.transatlantic-voice-archive
Transatlantic Voice Archive
Speech clips with aligned transcripts in the transatlantic (Mid-Atlantic)
accent — the clipped, semi-British delivery of 1930s–1960s American newsreel
announcers. Built from public-domain Universal
Newsreels (1929–1967) on the
Internet Archive, intended for finetuning TTS models on the accent.
Dataset statistics
Clips
1427
Total audio
2.04 h (122.2 min)
Average clip
5.14 s
Sample rate
22050 Hz, mono WAV
Source reels… See the full description on the dataset page: https://huggingface.co/datasets/brucemacd/transatlantic-voice-archive.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 2,116
Chunk audio duration: 6.624 hours
Source transcript rows represented: 5,601
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
56
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-01
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-01
This dataset contains lossless FLAC chunks derived from 45 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 1,034
Chunk audio duration: 4.411 hours
Source transcript rows represented: 3,041
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
9
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-01.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 1,097
Chunk audio duration: 5.143 hours
Source transcript rows represented: 3,875
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
19
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-04
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-04
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 599
Chunk audio duration: 4.567 hours
Source transcript rows represented: 2,961
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
0
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-04.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-05
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-05
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 1,536
Chunk audio duration: 5.773 hours
Source transcript rows represented: 4,326
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
22
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-05.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-03
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-03
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 650
Chunk audio duration: 5.893 hours
Source transcript rows represented: 4,213
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
1
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-03.
