datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ramanv-tts-all-raw
ramanv-tts-all-raw
Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages.
ellipsis-lrs3-raw
LRS3-TED — verified mirror
A mirror of the LRS3-TED dataset (Lip Reading Sentences 3), preserved because the official
distribution has been discontinued. This repository adds no new data: it is a re-hosted copy with a
full verification report against the official file list, so you know exactly what is and is not here.
Attribution
LRS3-TED was created by Triantafyllos Afouras, Joon Son Chung and Andrew Zisserman
(Visual Geometry Group, University of Oxford):
T.… See the full description on the dataset page: https://huggingface.co/datasets/TheNHz/ellipsis-lrs3-raw.northern-kurdish-raw-audio
Northern Kurdish Raw Audio Collection
Overview
This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources.
The collection was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-supervised Learning (SSL)
Spoken Language Understanding (SLU)
The dataset contains more than 2,000 hours… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-raw-audio.turkish-audiobook-raw
Turkish Audiobook Speech Corpus (Raw)
Türkçe konuşma araştırmaları için derlenmiş, işlenmemiş uzun-form ses kayıtlarından
oluşan bir koleksiyon. Kayıtlar çeşitli kaynaklardan bir araya getirilmiştir ve
konuşmacı, kayıt ortamı, süre ve ses kalitesi bakımından geniş bir çeşitlilik gösterir.
İçerik
Uzun-form Türkçe konuşma kayıtları (m4a / mp3)
Kaynağa göre klasörlenmiş düz dizin yapısı
Transkript, hizalama veya segmentasyon içermez — ham hâldedir… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-audiobook-raw.southern-kurdish-raw-audio
Southern Kurdish Raw Audio Collection
Overview
This repository contains approximately 170 hours of Southern Kurdish (SDH) raw speech collected from publicly available media sources, podcasts, interviews, news broadcasts, and online programs.
The main sources are Aryen TV and Kurd Channel.
The recordings mainly consist of spontaneous and semi-spontaneous speech, covering diverse speakers, topics, and acoustic conditions. While the majority of the content is… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/southern-kurdish-raw-audio.central-kurdish-audiobook-raw
Central Kurdish Audiobook Raw Audio Collection
Overview
This repository contains a large collection of raw Central Kurdish (Sorani Kurdish) audiobook recordings gathered from publicly available online sources.
The collection was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-supervised learning
The dataset contains approximately 4,300 hours of speech collected from 1026… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-audiobook-raw.raw-speech-whispervq-v1
Dataset Overview
This dataset contains over 2,4M English ASR samples, using:
The a training set of parler-tts/mls_eng_10k
Tokenized using WhisperVQ.
Usage
from datasets import load_dataset, Audio
# Load Instruction Speech dataset
dataset = load_dataset("homebrewltd/raw-speech-whispervq-v1",split='train')
Dataset Fields
Field
Type
Description
tokens
sequence
Tokenized using Encodec
text
sequence
Converted audio tokens
Bias, Risks… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/raw-speech-whispervq-v1.arabic-audio-collection-algerian-rawi
Rawi Postcast Arabic Speech Dataset
Dataset Summary
The Rawi Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 51 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-rawi.hawrami-kurdish-raw-audio
Hawrami Raw Audio Collection
Overview
This repository contains approximately 500 hours of Hawrami Kurdish raw speech collected from publicly available media sources.
The dataset was gathered primarily from the Rocyar program broadcast on Sterk TV, along with additional publicly available Hawrami-language content.
The recordings mainly consist of spontaneous and semi-spontaneous speech, including interviews, discussions, cultural programs, storytelling, and other… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/hawrami-kurdish-raw-audio.no-filter-raw-NepaliParliamentDSv2raw_1hr_myanmar_asr_audio
🇲🇲 Raw 1-Hour Burmese ASR Audio Dataset
A 1-hour dataset of Burmese (Myanmar language) spoken audio clips with transcripts, curated from official public-service media broadcasts by PVTV Myanmar — the media voice of Myanmar’s National Unity Government (NUG).
This dataset is intended for automatic speech recognition (ASR) and Burmese speech-processing research.
➡️ Author: freococo➡️ License: MIT➡️ Language: Burmese (my)
📦 Dataset Summary
Duration: ~1 hour
Chunks:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/raw_1hr_myanmar_asr_audio.quranic-asr-cloud-rawdata
Quranic ASR Provider Benchmark Results
Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark.
This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio.
What Is Included
Area
Path
Purpose
Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.2hr_myanmar_asr_raw_audio
🇲🇲 Raw 2-Hour Burmese ASR Audio Dataset
A ~2-hour Burmese (Myanmar language) ASR dataset featuring 1,612 audio clips with aligned transcripts, curated from official public-service educational broadcasts by FOEIM Academy — a civic media arm of FOEIM.ORG, operating under the Myanmar National Unity Government (NUG).
This dataset is MIT-licensed as a public good — a shared asset for the Burmese-speaking world. It serves speech technology, education, and cultural preservation efforts… See the full description on the dataset page: https://huggingface.co/datasets/freococo/2hr_myanmar_asr_raw_audio.tantraloka-dyczkowski-raw
Two views of the same 5,146 verses
structured is the full record — 24 columns over all 37 āhnikas, including extracted
entities, cross-references, parallel passages and audio alignment. Several of those
columns hold JSON documents running to thousands of characters, which is what makes the
dataset viewer unreadable in a browser: a row is a wall of text.
reading (the default config) is a projection of the same rows onto the eight columns a
reader wants — volume, chapter, verse… See the full description on the dataset page: https://huggingface.co/datasets/Anamavajra-Labs/tantraloka-dyczkowski-raw.kumu-livestream-raw
halo-livestream-raw
Unsegmented source recordings behind sapinsapin/halo-livestream — the archival input to the pipeline, not a training set.
1 recording(s) · 26:56 · 9.0MB
🔒 Gated on purpose
Full-length conversation between named, identifiable speakers is a very
different privacy proposition from the short segments in the processed
dataset, so access here is gated: request it and agree to the terms above.
If what you want is segmented, quality-scored… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/kumu-livestream-raw.3hr_myanmar_asr_raw_audio
📚 3-Hour Burmese Speech Dataset from FOEIM Academy (ASR-ready)
This is a curated ~3-hour dataset of Burmese-language audio-transcript pairs derived from the official public-service educational media of FOEIM Academy, a civic platform affiliated with FOEIM.ORG.
It is structured for fine-grained automatic speech recognition (ASR) training and testing.All data is aligned from timestamped subtitle files (.srt) and segmented into high-quality .mp3 mono files with aligned transcripts.
➡️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/3hr_myanmar_asr_raw_audio.goan-konkani-raw-audio
Goan Konkani Raw Audio
Private raw-ingestion index. Audio objects are preserved in their original YouTube audio container in the private Reubencf/goan-konkani-raw-audio Storage Bucket. Cleaning, repeated-Mass detection, song removal, segmentation, and transcription are derived stages; the raw source objects are not overwritten. Access and reuse remain subject to the source owners’ rights and permissions.
data/pilot_metadata.jsonl records checksums, durations, source URLs, and… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/goan-konkani-raw-audio.voa_9000h_raw_mp3
VOA 9000h Raw MP3
Raw MP3 archive of VOA Burmese radio broadcasts (2012–2025). 8,234 full-length programs, ~4,000 hours, ~138 GB.
Split across 17 tar files (voa_raw_0000.tar … voa_raw_0016.tar), 500 MP3s per tar.
Source: freococo/9000hours_voa_burmese_audio — filtered to live URLs.
Derived datasets:
voa_myanmar_voices — 20s FLAC chunks + transcripts (498 GB)
myanmar_asr — ASR model trained on this audio
License
Public domain (VOA staff recordings, U.S. 17 U.S.C.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_9000h_raw_mp3.nigerian-asr-dataset
NaijaVox ASR Dataset
An ASR dataset for four Nigerian languages — Hausa, Igbo, Yoruba, and
Nigerian Pidgin — sourced from Google's
Waxal corpus.
Built to train and evaluate automatic speech recognition models for
Nigerian languages.
Configs
Config
Language
Train
Validation
Test
Total
ha
Hausa
1655
296
20
1971
ig
Igbo
1604
287
20
1911
yo
Yoruba
2192
392
27
2611
pcm
Nigerian Pidgin
1674
299
20
1993
Splits: train 84% / validation 15% / test 1%… See the full description on the dataset page: https://huggingface.co/datasets/amn-raw/nigerian-asr-dataset.Podcast-Transcripts-Raw
Podcast Transcripts
Speaker-diarized transcripts from ~100k+ podcast episodes and YouTube videos (~45k hours of audio).
This dataset will be gated. Only people who are part of our team may access.
Splits
Config
Rows
Description
shows
124
Channels / podcast feeds (name, description, hosts, links)
episodes
102,374
Episode/video metadata (title, description, guests, tags, dates)
transcripts
102,374
ASR transcripts with diarized segments + YouTube… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Raw.raw-emocean
raw-emocean
Large-scale English speech dataset for text-to-speech (TTS) model training. Designed for autoregressive TTS architectures (TADA, CSM, VALL-E style models).
Dataset Summary
Metric
Value
Parquet shards
7
Segment duration
3–8 seconds
Sample rate
24,000 Hz (mono)
ASR engine
NVIDIA Parakeet TDT 0.6B v3
Format
Parquet with embedded audio
Dataset Schema
Column
Type
Description
audio
Audio
Waveform array + sampling… See the full description on the dataset page: https://huggingface.co/datasets/somu9/raw-emocean.
