datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stt-vibravox-fr-test
VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox)
Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour
benchmark ASR français multi-capteur sur audio standard ET non-standard
(bone-conduction, in-ear, throat, accéléromètre).
Ce repo contient uniquement les configs speech_clean + speech_noisy
(les seules avec transcription). Les configs speechless_* upstream sont
exclues car sans texte → pas de WER possible.
Configs
Config
Test rows
Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.fleurs_test
FLEURS Test Dataset with Enhanced Metadata
This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks.
Dataset Description
FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.Speech-MASSIVE-test
Speech-MASSIVE Test Split
This dataset repository is only for test split of Speech-MASSIVE.
train and dev splits are available in the separate dataset repository. https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE
Dataset Description
Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test.dhravani-mit-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-mit-test.RS-testdhravani-iitpatna-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitpatna-test.dhravani-IIT_Guwahati-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IIT_Guwahati-test.stt-cefc-fr-test
CEFC-Orfeo FR — long-form oral test mirror
Mirror non-officiel du Corpus d'Études du Français Contemporain (CEFC)
agrégé par le projet Orfeo (ANR), tel que distribué sur le portail
projet-orfeo.fr (release 13).
Long-form : 1 row = 1 fichier audio entier (30-60 min en moyenne).
12 sous-corpus oraux du français contemporain, 303 heures au total,
901 fichiers. Idéal pour bench Whisper / Canary en conditions réelles
(chunked decoding, dérive temporelle, multi-locuteurs).… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-cefc-fr-test.dhravani-IGDTUW_Delhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IGDTUW_Delhi-test.dhravani-iitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitdelhi-test.dhravani-iiitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iiitdelhi-test.edacc_teststt-covost25-test-fr
Common Voice FR — read + spontaneous + CoVoST 2 test mirror
Mirror combiné de trois ressources Mozilla / Facebook AI Research dans le
même repo pour benchmark WER/STT français multi-paradigme.
Config read (défaut)
Source upstream : Mozilla Common Voice 25.0 FR (paradigme lecture).
16 149 utterances, ~21 h. Locuteurs très variés (crowdsourced).
Paradigme : contributeurs lisent à voix haute des phrases écrites
d'un pool partagé. Débit régulier, peu d'hésitations.
Référence… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-covost25-test-fr.stt-summre-fr-test
SUMM-RE — French test split (mirror of linagora/SUMM-RE)
Mirror public du split test de SUMM-RE (LINAGORA / Aix-Marseille
LPL), pour benchmark ASR français conversationnel (parole de réunion,
3-4 locuteurs, ~20 min par session).
⚠ Ce repo ne contient que le split test (124 tracks individuelles =
37 réunions × 3-4 micros). Pour les splits train / dev, voir le repo
upstream linagora/SUMM-RE.
Contenu
124 pistes audio individuelles (1 piste = 1 microphone d'un locuteur… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-summre-fr-test.RS-test-fix2vi-asr-tech-test
Vietnamese ASR Test Set - Technology
A Vietnamese speech-recognition benchmark for the technology domain
(Công nghệ), released by G-Group AI Lab.
Audio is real-world Vietnamese speech covering consumer electronics reviews, software tutorials, programming and IT walkthroughs — dense in English loanwords and product names.
Listen & explore
Every utterance is playable inline in the viewer above — hit play on any row to
stream the clip. For full-text search across… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-tech-test.fastmss-v0.5.0-test
FastMSS synthetic multi-speaker meetings - parquet edition
Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring.
Subsets and splits
v0.5.0_test — splits: train — 1000 mixtures, 1081.0 min total, 3609… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/fastmss-v0.5.0-test.edacc_test_cleanvividh-test-hindi
🎙️ Vividh-ASR Benchmark — Hindi (Test Split)
How well does your ASR model actually work in the wild?
Vividh-ASR is a complexity-stratified benchmark that tells you exactly where your model succeeds — and where it falls apart.
Most Indic ASR benchmarks evaluate models on clean, studio-recorded speech. Real-world audio is not that. Vividh-ASR organises evaluation by acoustic complexity rather than domain, exposing the studio-bias that plagues models fine-tuned predominantly on read… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/vividh-test-hindi.ASR-GERMAN-MIXED-TEST
Dataset Beschreibung
Dieser Datensatz und die Beschreibung wurde von flozi00/asr-german-mixed übernommen und nur der Test-Split hier hochgeladen, da Hugging Face native erst einmal alle Splits herunterlädt. Für eine Evaluation von Speech-to-Text Modellen ist ein Download von 136 GB allerdings etwas zeit- & speicherraubend, weshalb wir hier nur den Test-Split für Evaluationen anbieten möchten. Die Arbeit und die Anerkennung sollten deshalb weiter bei primeline & flozi00 für die… See the full description on the dataset page: https://huggingface.co/datasets/avemio/ASR-GERMAN-MIXED-TEST.ami-2speaker-test
AMI 2-Speaker Test Set
Need a voice model for your domain? Trelis builds custom ASR, TTS, and voice agent pipelines for specialist verticals (legal, medical, finance, construction) and low-resource languages. Enquire or book a consultation →
A 50-clip benchmark for 2-speaker overlapping speech recognition, derived from the AMI Meeting Corpus test split.
Each clip is 8–28 seconds of real conversational meeting audio reconstructed as a 2-speaker virtual meeting, with separate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/ami-2speaker-test.vi-asr-finance-test
Vietnamese ASR Test Set - Finance
A Vietnamese speech-recognition benchmark for the finance domain
(Tài chính), released by G-Group AI Lab.
Audio is real-world Vietnamese speech covering stock market commentary, trading platforms, banking and crypto — dense in tickers, numbers and financial jargon.
Listen & explore
Every utterance is playable inline in the viewer above — hit play on any row to
stream the clip. For full-text search across transcripts, duration… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-finance-test.vi-asr-edu-test
Vietnamese ASR Test Set - Education
A Vietnamese speech-recognition benchmark for the education domain
(Giáo dục), released by G-Group AI Lab.
Audio is real-world Vietnamese speech covering study-abroad consulting, exam and certification guidance, university and training-course introductions.
Listen & explore
Every utterance is playable inline in the viewer above — hit play on any row to
stream the clip. For full-text search across transcripts, duration filters… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-edu-test.librispeech_asr_test_clean_word_timestamp
Word-level timestamp annotated Librispeech ASR test set
This dataset contains word-level timestamp information for the Librispeech ASR test (clean) dataset.
It contains 2620 short files that have been force-aligned with its text to get reasonably accurate word-level timestamp information.
Suitable for use in timestamp benchmarking of ASR models or audio dataset preprocessing.
To request access to more datasets like this, please fill out this form:… See the full description on the dataset page: https://huggingface.co/datasets/olympusmons/librispeech_asr_test_clean_word_timestamp.vi-asr-pubadmin-test
Vietnamese ASR Test Set - Public Administration
A Vietnamese speech-recognition benchmark for the public administration domain
(Hành chính công), released by G-Group AI Lab.
Audio is real-world Vietnamese speech covering administrative procedures, paperwork and licensing guidance, civil records — dense in place names and legal terminology.
Listen & explore
Every utterance is playable inline in the viewer above — hit play on any row to
stream the clip. For… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-pubadmin-test.cv10-uk-testset-clean
The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦
Overview
This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios.
All audios have been checked by a human to be sure that they are correct.
This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk
Community
Discord: https://bit.ly/discord-uds
Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.IWSLT2025-Test
Dataset details
This is the blind test set of IWSLT 2025's model compression track.
It consists of audio extracted from ACL presentations.
For training data in the same domain, the ACL 60/60 dataset can be used.
Citation
@INPROCEEDINGS{Abdulmumin2025-IWSLT,
title = "{Findings of the IWSLT 2025 Evaluation Campaign}",
author = "Abdulmumin, Idris and Agostinelli, Victor and Alumäe, Tanel and
Anastasopoulos, Antonios and {Ashwin} and Bentivogli… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/IWSLT2025-Test.vividh-test-malayalam
🎙️ Vividh-ASR Benchmark — Malayalam (Test Split)
How well does your ASR model actually work in the wild?Vividh-ASR is a complexity-stratified benchmark that tells you exactly where your model succeeds — and where it falls apart.
Most Indic ASR benchmarks evaluate models on clean, studio-recorded speech. Real-world audio is not that. Vividh-ASR organises evaluation by acoustic complexity rather than domain, exposing the studio-bias that plagues models fine-tuned predominantly on… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/vividh-test-malayalam.covost2-tr-test
covost2-tr-test
Turkish test split of CoVoST 2 (tr_en, Turkish source) (Common Voice–based ST/ASR corpus), re-hosted for Turkish STT benchmarking.
Rows: 1629
Columns: client_id, file, audio, sentence, translation, id (sentence = Turkish transcript, translation = English)
Source: https://github.com/facebookresearch/covost (audio mirror: fixie-ai/covost2)
License: cc0-1.0
Only the Turkish source test split is included, extracted as-is.
riksdagen_test
