datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/adab-tech/WaxalNLP.Treble10-Speech
Treble10-Speech (16 kHz)
The Treble10-Speech dataset is a dataset for automatic speech recognition (ASR), containing pre-convolved speech files using high fidelity room-acoustic simulations from the Treble10-RIR dataset with 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms.
The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s.
Examples:… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-Speech.quran-alignment-benchmark
Quran Recitation Alignment Benchmark
Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here.
16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.telugu-tech-custom-voice
🎙️ Telugu Tech Custom Voice Dataset
A high-quality, clean single-speaker Telugu Speech & Voice dataset tailored for training and fine-tuning neural Text-to-Speech (TTS) models (e.g. Coqui XTTS v2, Piper TTS, VITS, Bark) and Automatic Speech Recognition (ASR).
📊 Dataset Statistics
Total Clips: 455 audio files (.wav)
Total Audio Duration: 1 Hour 12 Minutes 48.5 Seconds (4,368.5 seconds)
Total Dataset Size: ~1.20 GB
Language: Telugu (te) with technical terms /… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-custom-voice.asr-vi
MCP Cloudwords VIVOS Processed ASR Dataset
This dataset contains Vietnamese speech data processed and prepared by MCP Cloudwords for ASR tasks.
It includes audio files and corresponding transcriptions, divided into train and test sets.
Replace YOUR_USERNAME/YOUR_DATASET_NAME with your actual Hugging Face username and dataset name
dataset = load_dataset("YOUR_USERNAME/YOUR_DATASET_NAME", trust_remote_code=True)
Display dataset information
print(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SMEW-TECH/asr-vi.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.Audio-Transcription-Models-Comparison-PT-BR
Audio Transcription Models Comparison
A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese.
About the Dataset
This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering:
Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.ne-en-codeswitching-asr-technical-interview
Dataset Summary
This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar.
It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.tech-podcast-audio-30s
Tech Podcast Audio 30s
This is a Japanese speech corpus derived from technology-related podcast audio distributed through publicly accessible podcast RSS feeds. Podcast episodes were segmented into clips of up to 29.9 seconds using voice activity detection.
The combined dataset contains 1,129,703 accepted clips totaling 5,516.14 hours (approximately 329 GiB of Parquet files). Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/tech-podcast-audio-30s.telugu-tech-indicf5-custom-voice
🎙️ Telugu Tech IndicF5 Custom Voice Dataset
A 100% verified, clean, single-speaker Telugu Speech & Voice dataset specially formatted and phonetically cleaned for training and fine-tuning ai4bharat/IndicF5 and neural Text-to-Speech (TTS) models.
All English technical terms, numbers, acronyms, and ASR mishearings have been converted into native Telugu phonetic script, cleaned of noise/brackets, and validated for optimal IndicF5 fine-tuning performance.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-indicf5-custom-voice.vi-asr-tech-test
Vietnamese ASR Test Set - Technology
A Vietnamese speech-recognition benchmark for the technology domain
(Công nghệ), released by G-Group AI Lab.
Audio is real-world Vietnamese speech covering consumer electronics reviews, software tutorials, programming and IT walkthroughs — dense in English loanwords and product names.
Listen & explore
Every utterance is playable inline in the viewer above — hit play on any row to
stream the clip. For full-text search across… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-tech-test.tech-speech-dataset
Tech Speech Dataset
A speech dataset built from tech tutorial YouTube videos (Techvangelists channel), chunked into ~30-second sentence-aware segments.
Dataset Details
337 audio clips of clean English speech
3 hours 18 minutes of audio
16kHz mono WAV format
Topics: AI, LLMs, Ollama, LangChain, web development, tech news
Columns
Column
Description
audio
Audio clip (16kHz mono)
text
Whisper transcript
duration
Clip duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/pianistprogrammer/tech-speech-dataset.commonvoice_vad_cy
Recordiadau Byr Common Voice Cymraeg
Mae'r set data yma wedi ei gynhyrchu drwy torri recordiadau maint arferol ac/neu hir o
Common Voice 18 , i darnau llai drwy defnyddio model
'Voice Activity Detection (VAD)'. Er enghraifft, gall recordiad o frawddeg fel "Ffotograffau a darluniau du-a-gwyn." cael ei thorri i dair
recordiad byr: "Ffotograffau", "a darluniau", "du-a-gwyn".
Defnyddwyd yn benodol split train_missing,
sef recordiadau o frawddegau set hyfforddi swyddogol Common… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/commonvoice_vad_cy.commonvoice_16_1_cy
Dataset Card for Welsh Common Voice Corpus 16.1
Dataset Details
Dataset Description
This dataset consists of all 114,139 MP3 recordings with corresponding text files from the Welsh language Common Voice 16.1 release.
It contains a total of 155.12 hours of speech from 1,832 contributors, 122.22 hours of which (or 78.79%) has been verified manually
by volunteers.
Mozilla's corpus creator command line tool was used, with the number of repeats
allowed for any… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/commonvoice_16_1_cy.Meta_STT_EN-IN_Tech_Interviews
Meta STT EN-IN Tech Interviews
Indian English speech recognition dataset sourced from technical interviews, annotated with rich speech metadata including age group, gender, emotion, and intent. Designed for training multi-task ASR models that jointly predict transcriptions and speaker attributes.
Dataset Details
Property
Value
Train examples
58,000
Validation examples
1,204
Language
English (Indian accent)
Audio
16 kHz
Total size
~28 GB… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN-IN_Tech_Interviews.commonvoice_18_0_cy
Common Voice Cymraeg (fersiwn 18, Mehefin 2024)
Mae'r set ddata hon yn cynnwys holl recordiadau MP3 gyda ffeiliau testun cyfatebol o Common Voice 18.0 Cymraeg wedi eu drefnu i
wahanol is-casgliadau ('splits').
Defnyddwyd CorpusCreator gan Mozilla gyda'r nifer o ailrecordiadau o frawddegau
a chaniateir yn 99 (s=99). Nid yw brawddegau'r tri split -s99 yn cydanghynhwysol. (h.y. gall brawddeg bodoli mewn mwy nag un split)
Manylion pob split
Split
Duration
Word… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/commonvoice_18_0_cy.banc-trawsgrifiadau-bangor_mt
Dataset Card for banc-trawsgrifiadau-bangor_mt
Source Dataset Provenance
This dataset was derived from the source dataset at commit SHA: da64c66e467dac267ee3aeaa2f520ef82722379e
(short: da64c66e)
commonvoice_tts_cy
Corpws TTS Common Voice Cymraeg
Mae'r set data hon yn cynnwys allbwn testun-i-leferydd Cymraeg techiaith o frawddegau sydd heb eu recordio o fewn
Commonvoice Cymraeg fersiwn 18.
Manylion pob split
Split
Duration
Word Count
No Of Clips
train
78:50:12
685,853
69,812
test
00:47:42
6896
706
commonvoice_18_0_cy_en
Common Voice Ddwyieithog (Cymraeg a Saesneg)
Mae'r set data hon yn cynnwys cyfuniad cyfartal o recordiadau Cymraeg a Saesneg o Common Voice fersiwn 18.
O Common Voice Cymraeg defnyddwyd splitau
train_all ac
other_with_excluded
ar gyfer recordiadau Cymraeg set hyfforddi'r set data hon. Cymerwyd yr un nifer o recordiadau o set hyfforddi swyddogol
Saesneg Common Voice fersiwn 18, gan blaenoriaethu'r rhai wedi eu dagio â acen Saesneg o Ynysoedd Prydain.
(h.y. Cymreig, Albanaidd… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/commonvoice_18_0_cy_en.
