datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ami
Dataset Card for AMI
Dataset Description
The AMI Meeting Corpus consists of 100 hours of meeting recordings. The recordings use a range of signals
synchronized to a common timeline. These include close-talking and far-field microphones, individual and
room-view video cameras, and output from a slide projector and an electronic whiteboard. During the meetings,
the participants also have unsynchronized pens available to them that record what is written. The meetings
were… See the full description on the dataset page: https://huggingface.co/datasets/edinburghcstr/ami.edacc
EdAcc: The Edinburgh International Accents of English Corpus
The Edinburgh International Accents of English Corpus (EdAcc) is a new automatic speech recognition (ASR) dataset
composed of 40 hours of English dyadic conversations between speakers with a diverse set of accents. EdAcc includes a
wide range of first and second-language varieties of English and a linguistic background profile of each speaker.
Results on latest public, and commercial models show that EdAcc highlights… See the full description on the dataset page: https://huggingface.co/datasets/edinburghcstr/edacc.vctkThe CSTR VCTK Corpus includes speech data uttered by 110 English speakers with various accents.voice-acting-edge-top3
Edge-case reward-Top-3 — voice annotations
1,494,503 synthetic expressive-speech utterances (the reward-Top-3 selection of the
edge-case corpus), annotated with:
a CrisperWhisper-format transcript with corrected vocal bursts — each surviving burst
carries a class name from laion/vocal-burst-detector-v2
at its original timestamp;
raw voice scores at two granularities — whole utterance and per sentence — from
laion/Empathic-Insight-Voice-Plus
(all 40 emotions), the 57 VoiceNet… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-edge-top3.romanian-speech-v2
Research Use Only — This dataset is released strictly for personal research and educational
purposes. The processing pipeline and all scripts are fully open source, but the underlying audio
originates from sources with varying copyrights. Only short fragments were used under fair use
provisions and EU Copyright Directive Art. 3 (text and data mining for scientific research).
This dataset must not be used for redistribution of the source material, commercial purposes,
or training commercially… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-speech-v2.hyvoxpopuli
HyVoxPopuli
HyVoxPopuli is an open Armenian speech dataset (~6 hours, 16 kHz mono) with raw and normalized transcripts. It targets ASR and TTS research for Eastern Armenian (hy-AM).
Note: Despite the name, this release is not the official Facebook VoxPopuli parliament corpus. Audio is literary narration (two voice actors) segmented into short clips. Update citations and experiments accordingly.
Dataset summary
Rows
623
Train / Val / Test
498 / 62… See the full description on the dataset page: https://huggingface.co/datasets/Edmon02/hyvoxpopuli.eduskunta-asr
Finnish Parliament ASR Dataset
Speech dataset from Finnish parliament (Eduskunta) plenary sessions with aligned ASR transcriptions and human reference texts.
Overview
Train
Test
Segments
64,945
9,121
Duration
423 h
59 h
Speakers
212
187
Total: 74,066 segments, 482 hours, 212 unique speakers.
Audio is 16 kHz mono Opus. Segments are individual speaker turns. Segments with CER > 0.30 are excluded (see below).
Schema
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/samuli/eduskunta-asr.ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.french-education-speech
French Education Speech - Transcribed Dataset
High-quality French educational speech dataset transcribed with OpenAI Whisper API, prepared for training automatic speech recognition (ASR) models.
Dataset Summary
This dataset contains 3,933 transcribed audio segments from the French educational domain, totaling approximately 12.82 hours of audio. All transcriptions were performed using OpenAI Whisper API (optimized Whisper-1 model) to ensure maximum accuracy, especially for… See the full description on the dataset page: https://huggingface.co/datasets/MEscriva/french-education-speech.massive-yt-edu-queue
Massive YouTube Educational Video Queue
Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours.
Description
This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.uzbek-asr-curated-701h
Uzbek ASR Curated Dataset (701 hours)
A curated multi-source Uzbek speech dataset for automatic speech recognition (ASR) training and evaluation.
Dataset Description
Language
Uzbek (Latin script with okina ʻ)
Total utterances
337,920
Total duration
~701 hours
Audio format
16 kHz mono WAV (PCM_16)
Manifest format
NeMo JSONL
Splits
train (94%) / val (3%) / test (3%)
Splits
Split
Utterances
Hours
Train
317,655… See the full description on the dataset page: https://huggingface.co/datasets/uzinfocom-edu-ai/uzbek-asr-curated-701h.edacc_testRTPSpeechPortuguese Speech
This dataset aims to provide people with European Portuguese audio and textual data pairs, which can be used to fine-tune large language models.
These Portuguese recordings are from RTP (Rádio e Televisão de Portugal), which we have broken down and transcribed into short sentences.
Citation
If you use this dataset, please cite:
L. M. Hoi, Y. Sun and S. K. Im, "An Automatic Speech Segmentation Algorithm of Portuguese based on Spectrogram Windowing," 2022 IEEE World AI IoT… See the full description on the dataset page: https://huggingface.co/datasets/edmond5995/RTPSpeech.romanian-tts-single-speaker
Romanian TTS Single Speaker
A single-speaker Romanian speech dataset for TTS model training.
Dataset Description
Segments
24,379
Duration
34.3 hours
Speaker
Sanda (female)
Language
Romanian (ro)
Audio
WAV, 16-bit, mono, 24 kHz
Subsets
Subset
Segments
Description
standard
24,203
Standard Romanian sentences
loanword
176
Sentences containing foreign loanwords
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-tts-single-speaker.edacc_test_cleanmassive-yt-edu-transcriptions
Massive YouTube Educational Transcriptions
Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5.
Stats
Videos: 59,355
Characters: 1,539,022,925 (~384M tokens)
Audio hours: 35,890
Model: faster-whisper (CTranslate2) with distil-large-v3.5
Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime
Fields
Field
Description
video_id
YouTube video ID
title
Video title
text
Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.french-education-speech
French Education Speech - Transcribed Dataset
High-quality French educational speech dataset transcribed with OpenAI Whisper API, prepared for training automatic speech recognition (ASR) models.
Dataset Summary
This dataset contains 3,933 transcribed audio segments from the French educational domain, totaling approximately 12.82 hours of audio. All transcriptions were performed using OpenAI Whisper API (optimized Whisper-1 model) to ensure maximum accuracy, especially for… See the full description on the dataset page: https://huggingface.co/datasets/Lexia-Labs/french-education-speech.vi-asr-edu-test
Vietnamese ASR Test Set - Education
A Vietnamese speech-recognition benchmark for the education domain
(Giáo dục), released by G-Group AI Lab.
Audio is real-world Vietnamese speech covering study-abroad consulting, exam and certification guidance, university and training-course introductions.
Listen & explore
Every utterance is playable inline in the viewer above — hit play on any row to
stream the clip. For full-text search across transcripts, duration filters… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-edu-test.edu-subject
Education Subject Samples
Note: The audio samples presented here have been compressed to MP3 for browser playback and do not reflect the actual acoustic quality of the dataset. Refer to the Technical Specs below for the specifications of what will actually be delivered.
Education Subject Samples is a speech dataset of real tutoring conversations between tutors and people seeking help on various academic subjects. The dataset is captured in real-life settings, covering a… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/edu-subject.EdgeMMEval
EdgeMMEval
Minimal multimodal evaluation dataset for on-device inference testing.
Covers functional correctness, accuracy, latency stress, and memory
pressure across image, audio, text, multi-turn, combination, structured
output, and tool-calling cases.
Dataset summary
The test split is defined in data/test/metadata.jsonl (200 rows). Each
row has a test_id (for example IMG-001, STO-020) and a modality.
Modality
Samples
Focus
Image
34
VQA, OCR, description… See the full description on the dataset page: https://huggingface.co/datasets/CortexSwarm/EdgeMMEval.educatinayt
Dataset Card for "educatinayt"
More Information needed
GroupG_Project
KPL Esports Linguistic Dataset
Abstract
This dataset introduces a multimodal labeled dataset including 65 minutes and 24 seconds of commentary from three dynamic matches in the King Pro League (KPL) of the game "Honor of Kings": the 2021 Autumn Semifinals, 2025 Finals, and 2025 Summer Finals. This dataset was collected by extracting AI-generated, time-aligned transcripts from recording videos on BiliBili Platform, which combines essential manually modification to ensure… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/GroupG_Project.droidnexus-arabic-editorial-speech-scorecard-mini
DroidNexus Arabic Editorial Speech Scorecard Mini
A public DroidNexus Labs scorecard dataset for Arabic speech workflows: representative editorial scenarios, latency targets, overlap pressure, and the metric stack that decides whether a transcript is usable.
Why this exists
This dataset is the first public speech artifact layer for DroidNexus Labs. It publishes representative editorial workloads and evaluation pressure before claiming a full source-audio benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/droidnexus-arabic-editorial-speech-scorecard-mini.new_taipei_education_department_dataset
海岸阿美語語音資料集
資料集摘要
本資料集收錄台灣海岸阿美語的語音與文字對照,適用於自動語音辨識(ASR)、語音合成(TTS)及原住民族語相關研究。
每筆資料包含一段語音、羅馬拼音轉寫、中文翻譯,以及語別標籤。
項目
說明
語言
海岸阿美語(Amis)
樣本數
33,080
Split
train
音檔格式
MP3
平均時長
約 6.6 秒
時長範圍
1.0 – 117.5 秒
資料欄位
欄位
型別
說明
範例
id
string
唯一識別碼
record1-1@0001.mp3
audio
audio
語音檔
(可於 Dataset Viewer 播放)
duration
float64
音檔長度(秒)
3.288
transcript
string
羅馬拼音轉寫
cecay
translation
string
中文翻譯
一
lang_group
string
語別(中文)
海岸阿美語… See the full description on the dataset page: https://huggingface.co/datasets/allen-1216/new_taipei_education_department_dataset.Gujarati40heducation_chunked
education_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/education_chunked.education_chunked_speech_restorised
education_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/education_chunked_speech_restorised.educational_unchunked
educational_unchunked
This is a gated Uzbek unchunked source speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/educational_unchunked.education_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/education_chunked
Aligned dataset: instinct-org/education_chunked_nfa_aligned
Rows: 386961 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/education_chunked_nfa_aligned.
