datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
speechio_test
SpeechIO ASR Test Sets (parquet)
Parquet repackaging of the SpeechColab SpeechIO Mandarin ASR benchmark,
re-exported from yuekai/speechio (Lhotse cuts) into standard
HuggingFace parquet with embedded 16 kHz audio.
27 test sets: SPEECHIO_ASR_ZH00000 ... SPEECHIO_ASR_ZH00026, each a config with a single test split.
~43k utterances, ~66 hours total, evaluation only.
Columns
column
type
note
segment_id
string
utterance id
speaker
string
speaker id… See the full description on the dataset page: https://huggingface.co/datasets/yuekai/speechio_test.fleurs_test
FLEURS Test Dataset with Enhanced Metadata
This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks.
Dataset Description
FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.testThe Snow Mountain dataset contains the audio recordings (in .mp3 format) and the corresponding text of The Bible
in 11 Indian languages. The recordings were done in a studio setting by native speakers. Each language has a single
speaker in the dataset. Most of these languages are geographically concentrated in the Northern part of India around
the state of Himachal Pradesh. Being related to Hindi they all use the Devanagari script for transcription.Speech-MASSIVE-test
Speech-MASSIVE Test Split
This dataset repository is only for test split of Speech-MASSIVE.
train and dev splits are available in the separate dataset repository. https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE
Dataset Description
Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test.RS-teststt-vibravox-fr-test
VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox)
Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour
benchmark ASR français multi-capteur sur audio standard ET non-standard
(bone-conduction, in-ear, throat, accéléromètre).
Ce repo contient uniquement les configs speech_clean + speech_noisy
(les seules avec transcription). Les configs speechless_* upstream sont
exclues car sans texte → pas de WER possible.
Configs
Config
Test rows
Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.edacc_testedacc_test_cleandhravani-mit-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-mit-test.ciempiess_testThe CIEMPIESS TEST Corpus is a gender balanced corpus destined to test acoustic models for the speech recognition task. The corpus was manually transcribed and it contains audio recordings from 10 male and 10 female speakers. The CIEMPIESS TEST is one of the three corpora included at the LDC's \"CIEMPIESS Experimentation\" (LDC2019S07).dhravani-iitpatna-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitpatna-test.dhravani-IIT_Guwahati-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IIT_Guwahati-test.stt_test_audio
Mobile Voice Platform STT Test Audio
Versioned benchmark audio for the MobileVoicePlatform-Android SampleApp. The repository intentionally
contains two views of the same clips:
data/ is an AudioFolder-compatible view with metadata.csv for Hugging Face tooling.
packs/ contains checksum-pinned ZIPs optimized for bounded download and validation on Android.
catalog.json is the machine-readable index used to discover datasets, categories, languages,
checksums, clip references, and… See the full description on the dataset page: https://huggingface.co/datasets/sfsmcnulty/stt_test_audio.dhravani-IGDTUW_Delhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IGDTUW_Delhi-test.dhravani-iitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitdelhi-test.dhravani-iiitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iiitdelhi-test.audio-testing
audio-testing
Overview
This is a small, open dataset designed for quick validation of audio-related pipelines and applications, especially for Text-to-Speech (TTS) and Speech-to-Text (STT) systems.
It provides a few short, diverse audio clips and corresponding text transcripts, allowing developers to verify input/output handling, audio processing, and transcription logic without downloading large datasets.
Contents
3 short audio samples (.mp3, .wav)… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/audio-testing.stt-covost25-test-fr
Common Voice FR — read + spontaneous + CoVoST 2 test mirror
Mirror combiné de trois ressources Mozilla / Facebook AI Research dans le
même repo pour benchmark WER/STT français multi-paradigme.
Config read (défaut)
Source upstream : Mozilla Common Voice 25.0 FR (paradigme lecture).
16 149 utterances, ~21 h. Locuteurs très variés (crowdsourced).
Paradigme : contributeurs lisent à voix haute des phrases écrites
d'un pool partagé. Débit régulier, peu d'hésitations.
Référence… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-covost25-test-fr.h-test@misc{humair025/h-test,
title = {h-test},
author = {Humair Munir},
year = {2025},
howpublished = {\url{https://huggingface.co/datasets/humair025/h-test}},
note = {Synthetic dataset of speech (with emotions). Licensed under CC-BY 4.0.}
}
RS-test-fix2fastmss-v0.5.0-test
FastMSS synthetic multi-speaker meetings - parquet edition
Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring.
Subsets and splits
v0.5.0_test — splits: train — 1000 mixtures, 1081.0 min total, 3609… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/fastmss-v0.5.0-test.vi-asr-tech-test
Vietnamese ASR Test Set - Technology
A Vietnamese speech-recognition benchmark for the technology domain
(Công nghệ), released by G-Group AI Lab.
Audio is real-world Vietnamese speech covering consumer electronics reviews, software tutorials, programming and IT walkthroughs — dense in English loanwords and product names.
Listen & explore
Every utterance is playable inline in the viewer above — hit play on any row to
stream the clip. For full-text search across… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-tech-test.X-Voice-TestsetX-Voice Multilingual Test Set
High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages
This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model.
Dataset Summary
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.test4
test4
This is a merged speech dataset containing 345 audio segments from 2 source datasets.
Dataset Information
Total Segments: 345
Speakers: 7
Languages: en
Emotions: neutral, sad, angry, happy
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion (neutral, happy… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test4.vi-asr-finance-test
Vietnamese ASR Test Set - Finance
A Vietnamese speech-recognition benchmark for the finance domain
(Tài chính), released by G-Group AI Lab.
Audio is real-world Vietnamese speech covering stock market commentary, trading platforms, banking and crypto — dense in tickers, numbers and financial jargon.
Listen & explore
Every utterance is playable inline in the viewer above — hit play on any row to
stream the clip. For full-text search across transcripts, duration… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-finance-test.ASR-GERMAN-MIXED-TEST
Dataset Beschreibung
Dieser Datensatz und die Beschreibung wurde von flozi00/asr-german-mixed übernommen und nur der Test-Split hier hochgeladen, da Hugging Face native erst einmal alle Splits herunterlädt. Für eine Evaluation von Speech-to-Text Modellen ist ein Download von 136 GB allerdings etwas zeit- & speicherraubend, weshalb wir hier nur den Test-Split für Evaluationen anbieten möchten. Die Arbeit und die Anerkennung sollten deshalb weiter bei primeline & flozi00 für die… See the full description on the dataset page: https://huggingface.co/datasets/avemio/ASR-GERMAN-MIXED-TEST.ami-2speaker-test
AMI 2-Speaker Test Set
Need a voice model for your domain? Trelis builds custom ASR, TTS, and voice agent pipelines for specialist verticals (legal, medical, finance, construction) and low-resource languages. Enquire or book a consultation →
A 50-clip benchmark for 2-speaker overlapping speech recognition, derived from the AMI Meeting Corpus test split.
Each clip is 8–28 seconds of real conversational meeting audio reconstructed as a 2-speaker virtual meeting, with separate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/ami-2speaker-test.vi-asr-edu-test
Vietnamese ASR Test Set - Education
A Vietnamese speech-recognition benchmark for the education domain
(Giáo dục), released by G-Group AI Lab.
Audio is real-world Vietnamese speech covering study-abroad consulting, exam and certification guidance, university and training-course introductions.
Listen & explore
Every utterance is playable inline in the viewer above — hit play on any row to
stream the clip. For full-text search across transcripts, duration filters… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-edu-test.vividh-test-hindi
🎙️ Vividh-ASR Benchmark — Hindi (Test Split)
How well does your ASR model actually work in the wild?
Vividh-ASR is a complexity-stratified benchmark that tells you exactly where your model succeeds — and where it falls apart.
Most Indic ASR benchmarks evaluate models on clean, studio-recorded speech. Real-world audio is not that. Vividh-ASR organises evaluation by acoustic complexity rather than domain, exposing the studio-bias that plagues models fine-tuned predominantly on read… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/vividh-test-hindi.vi-asr-pubadmin-test
Vietnamese ASR Test Set - Public Administration
A Vietnamese speech-recognition benchmark for the public administration domain
(Hành chính công), released by G-Group AI Lab.
Audio is real-world Vietnamese speech covering administrative procedures, paperwork and licensing guidance, civil records — dense in place names and legal terminology.
Listen & explore
Every utterance is playable inline in the viewer above — hit play on any row to
stream the clip. For… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-pubadmin-test.
