datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
riigikogu-audio-stenograms-2018-2025
Riigikogu Stenograms 2018-2025
This dataset contains stenograms (approximate transcripts) of the sessions of the Estonian Parliament Riigikogu, spanning a time period from the very end of 2017 to May 2025, together with the corresponding audio.
The transcripts are not verbatim (word-by-word) transcripts but are edited for readability and grammatical correctness. Sentence start and end times (w.r.t. to the corresponding audio file) are provided.
The transcripts are converted into… See the full description on the dataset page: https://huggingface.co/datasets/TalTechNLP/riigikogu-audio-stenograms-2018-2025.uts2025_vietipa
Vietnamese IPA Dataset
A comprehensive Vietnamese IPA (International Phonetic Alphabet) dataset with word pronunciations and MP3 audio files for text-to-speech and pronunciation learning applications.
Dataset Description
Dataset Summary
This dataset contains 50 common Vietnamese words with their IPA (International Phonetic Alphabet) transcriptions and corresponding audio files. It's designed for:
Text-to-speech systems development
Vietnamese pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/uts2025_vietipa.IWSLT2025-Test
Dataset details
This is the blind test set of IWSLT 2025's model compression track.
It consists of audio extracted from ACL presentations.
For training data in the same domain, the ACL 60/60 dataset can be used.
Citation
@INPROCEEDINGS{Abdulmumin2025-IWSLT,
title = "{Findings of the IWSLT 2025 Evaluation Campaign}",
author = "Abdulmumin, Idris and Agostinelli, Victor and Alumäe, Tanel and
Anastasopoulos, Antonios and {Ashwin} and Bentivogli… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/IWSLT2025-Test.data-kit-sub-iwslt2025-if-long-constraint
Data for KIT’s Instruction Following Submission for IWSLT 2025
This repo contains the data used to train our model for IWSLT 2025's Instruction-Following (IF) Speech Processing track.
IWSLT 2025's Instruction-Following (IF) Speech Processing track in the scientific domain aims to benchmark foundation models that can follow natural
language instructions—an ability well-established in textbased LLMs but still emerging in speech-based counterparts. Our approach employs an end-to-end… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/data-kit-sub-iwslt2025-if-long-constraint.2025-zwesui-g02-medyczna
ZWESUI 2025 - Grupa 2 - medyczna (mowa syntetyczna)
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2025 (pierwsza), tryb niestacjonarny.
Zespol (atrybucja): Grupa 2 (2025)
Zrodlo oryginalne: https://huggingface.co/datasets/yanvoi/med_male_female_r2
Domena: medyczna
Opis: 200 zdań z terminologią medyczną, mowa syntetyczna (ElevenLabs), głosy męskie i żeńskie; zbiór użyty… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2025-zwesui-g02-medyczna.medical-terms-2025
Medical Terms 2025 — Medical ASR Benchmark
Entity-aware medical ASR benchmark — 50 hard rows with synthetic TTS audio of 2025 drug/condition terminology.
Prepared by Trelis Research. Watch more on Youtube or inquire about our custom voice AI (ASR/TTS) services here.
Source
84 manually curated terms from 2025 FDA/EMA/WHO primary sources. Each term has source_url, source_date, and source_quality. Sentences generated by Gemini 2.5 Flash. Audio by Kokoro TTS via Trelis Studio… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/medical-terms-2025.2025-zwesui-g03-debata-prezydencka
ZWESUI 2025 - Grupa 3 - debata prezydencka 2025
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2025 (pierwsza), tryb niestacjonarny.
Zespol (atrybucja): Grupa 3 (2025)
Zrodlo oryginalne: https://huggingface.co/datasets/directtt/polish_presidential_debate
Domena: debata prezydencka
Opis: Debata prezydencka TVP z 12 maja 2025 — 13 kandydatów, 195 wypowiedzi, mowa spontaniczna.… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2025-zwesui-g03-debata-prezydencka.dataset-20250728_102101-sw
GeoPoll Swahili Speech Dataset
This dataset contains speech recognition data for Swahili (sw) collected and processed by GeoPoll.
Dataset Summary
This dataset is designed for fine-tuning speech recognition models on Swahili audio data. It includes high-quality audio segments with corresponding transcriptions.
Dataset Statistics
Total samples: 11814
Total duration: 20.45 hours
Average duration: 6.23 seconds per sample
Number of speakers: 6
Language: Swahili… See the full description on the dataset page: https://huggingface.co/datasets/GeoPoll/dataset-20250728_102101-sw.
