datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rixvox-v2
RixVox-v2: A Swedish parliamentary speech dataset
RixVox-v2 is a parliamentary speech dataset spanning nearly 23000 hours of speech. The dataset was built by matching and force aligning speeches in parliamentary protocols to media recordings of debates. Each observation contains metadata about the speaker's name, gender, district, role, party affiliation, and the date the speech was given. We include identifiers for protocols, speeches and speakers that allow linking observations in… See the full description on the dataset page: https://huggingface.co/datasets/KBLab/rixvox-v2.cv_corpus_v22
Dataset Card for Common Voice Corpus 22.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
NOTE: currently converting to parquet for convenience.. WIP
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.japanese-anime-speech-v2
Japanese Anime Speech Dataset V2
日本語はこちら
japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models.
The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels.
This dataset is not an updated version of japanese-anime-speech-v1.
For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset.
The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2.dialectal-arabic-lahgtna-v2
Dialectal Arabic Lahgtna v2
Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI.
Dataset Summary
~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech
**13 Arabic dialects **, labeled per utterance
16 kHz mono audio
Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.yt-danish-public-v2GLOBE_V2
Important notice
Differences between V2 version and the version described in paper:
The V2 version provide audio in 44.1kHz sample rate. (Supersampling)
The V2 versionn removed some samples (~5%) due to the volumn and text aligment issues.
Globe
The full paper can be accessed here: arXiv
An online demo can be accessed here: Github
Abstract
This paper introduces GLOBE, a high-quality English corpus with worldwide accents, specifically designed to address the… See the full description on the dataset page: https://huggingface.co/datasets/MushanW/GLOBE_V2.reazon-speech-v2-clone
Reazon Speech v2 dataset mirror
Original Dataset Source
Hugging Face Dataset Page: reazon-research/reazonspeech
Project Page: Reazon Research
License
This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction:
TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.ivirits-audio-v2-30s
ivrit.ai audio-v2 — 2–30 s segments
ivrit-ai/audio-v2 (>20k hours of Hebrew
audio) cut into 2–30 second speech segments with machine transcripts, ready for ASR
fine-tuning.
How it was built
VAD — Silero VAD (ONNX) over each episode decoded to 16 kHz mono. Speech regions
longer than 30 s are split at the quietest sufficiently-long pause inside the window,
so cuts land in silence rather than mid-word. Regions shorter than 2 s are dropped.
Transcription —… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/ivirits-audio-v2-30s.Japanese-Eroge-Voice-V2
Japanese-Eroge-Voice-V2
Description
This is the successor to the Japanese-Eroge-Voice dataset. It consists of a significantly larger collection of audio-transcription pairs extracted from Japanese eroge (adult games).
Note on Versioning: There is no overlap between this dataset (V2) and the previous version. All audio clips and transcriptions in V2 are distinct from those in the original version, providing entirely new data for research.
This version (V2) expands the… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice-V2.japanese-anime-speech-v2
Japanese Anime Speech Dataset V2
日本語はこちら
japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models.
The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels.
This dataset is not an updated version of japanese-anime-speech-v1.
For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset.
The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/WatsonNT/japanese-anime-speech-v2.romanian-speech-v2
Research Use Only — This dataset is released strictly for personal research and educational
purposes. The processing pipeline and all scripts are fully open source, but the underlying audio
originates from sources with varying copyrights. Only short fragments were used under fair use
provisions and EU Copyright Directive Art. 3 (text and data mining for scientific research).
This dataset must not be used for redistribution of the source material, commercial purposes,
or training commercially… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-speech-v2.japanese-anime-speech-v2-split-150k
japanese-anime-speech-v2-split-150k
joujiboi/japanese-anime-speech-v2 的 150,000 筆隨機子集,已切好 train / test。
split
rows
train
135,000
test
15,000
total
150,000
Columns
audio — 16 kHz mp3,與原始資料完全相同(未重新編碼)
sentence — 轉錄文字(原始欄位名為 transcription)
How it was built
來源的 sfw(271,788 筆)與 nsfw(20,849 筆)兩個 split 都有使用,並依原始比例分配名額
(sfw 139,313 / nsfw 10,687)。
在每個 split 的全部列上做無放回均勻抽樣,因此每筆資料被選中的機率相同。
抽出後整體打亂,前 15,000 筆為 test,其餘為 train,train / test… See the full description on the dataset page: https://huggingface.co/datasets/hhim8826/japanese-anime-speech-v2-split-150k.composite_corpus_eu_v2.1
Composite dataset for Basque made from public available data
This dataset is composed of the following public available data:
Train split:
The train split is composed of the following datasets combined:
mozilla-foundation/common_voice_18_0/eu: "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data)
gttsehu/basque_parliament_1/eu: "train_clean" split removing some of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_eu_v2.1.afrivox-v2
AfriVox-v2: A Domain-Verticalized Benchmark for In-the-Wild African Speech Recognition
AfriVox-v2 is a comprehensive multilingual speech recognition benchmark designed to evaluate ASR systems under realistic African deployment conditions. It covers 24 African languages across 10 application domains, with a strong emphasis on spontaneous, unscripted "in the wild" audio.
Dataset Summary
Most existing ASR benchmarks for African languages rely on scripted, read… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/afrivox-v2.coral-v2
CoRal: Danish Conversational and Read-aloud Dataset
Dataset Overview
CoRal is a comprehensive Automatic Speech Recognition (ASR) dataset designed to capture the diversity of the Danish language across various dialects, accents, genders, and age groups. The primary goal of the CoRal dataset is to provide a robust resource for training and evaluating ASR models that can understand and transcribe spoken Danish in all its variations.
Key Features
Dialect and… See the full description on the dataset page: https://huggingface.co/datasets/CoRal-project/coral-v2.BrahuiSpeech-70H-V2
BrahuiSpeech-70H V2
BrahuiSpeech-70H V2 is an approximately 70-hour automatic speech recognition dataset containing 15,626 audio-transcription pairs and 69 hours, 32 minutes, 12 seconds of real-world Brahui (Brahvi) speech.
Brahui (brh) is a low-resource Dravidian language spoken primarily in Balochistan, Pakistan. The dataset covers naturally occurring speech across varied speakers, speaking styles, media domains, and acoustic conditions. Transcriptions use the Perso-Arabic… See the full description on the dataset page: https://huggingface.co/datasets/TBOGamer22/BrahuiSpeech-70H-V2.Tamazight-ASR-Dataset-v2
Tamazight-Arabic Speech Recognition Dataset
Dataset Description
This dataset contains speech segments in Tamazight (specifically focusing on the Tachelhit dialect) paired with their corresponding Arabic transcriptions. It is designed to support the development of automatic speech recognition (ASR) systems for the Tamazight language, particularly for translation into Modern Standard Arabic.
This is an actively growing dataset, with regular updates and new data points being… See the full description on the dataset page: https://huggingface.co/datasets/SoufianeDahimi/Tamazight-ASR-Dataset-v2.audios-lingala-annotatees-v2
Annotated Lingala Audio — canonical corpus
Annotated Lingala speech for open automatic speech recognition research and for
fine-tuning speech models.
This release is a full reconstruction of the corpus from its source
recordings and annotations. It supersedes
Congo-digital-service/audios-lingala-annotatees,
which is deprecated — see Relationship to the previous release below.
What this dataset contains
Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.large-dataset-audio-v2
khmer_speech_dataset
Khmer speech dataset with transcriptions, speaker labels, and metadata.
Dataset Description
This dataset contains Khmer (Cambodian) speech recordings with detailed transcriptions and annotations.
Dataset Statistics
Metric
Value
Total Examples
9,285
Total Duration
336.68 hours
Average Duration
130.54 seconds
Total Words
1,336,611
Khmer Words
994,050 (74.4%)
English Words
340,127 (25.4%)
Unique Sources
2857
Unique… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/large-dataset-audio-v2.pl-asr-bigos-v2BIGOS (Benchmark Intended Grouping of Open Speech) dataset goal is to simplify access to the openly available Polish speech corpora and
enable systematic benchmarking of open and commercial Polish ASR systems.fleurs-ethiopian-v2
FLEURS — Ethiopian Languages
This dataset is a restructured v2 conversion of the Google FLEURS dataset for two Ethiopian languages: Amharic (am_et) and Oromo (om_et).
Subsets
Subset
Language
ISO 639-2
Train
Dev
Test
amh
Amharic
amh
3,163
223
516
orm
Oromo
orm
1,701
19
41
Splits
Split
Description
train
Training split
dev
Development split (renamed from validation in original FLEURS)
test
Test split
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/fleurs-ethiopian-v2.audio-3-speaker-dataset-v2
Three-Speaker Audio Dataset with Timbral Speaker Embeddings
A teaching dataset maintained by AI-Academy. It pairs single-speaker English
utterances with precomputed timbral speaker embeddings, and is intended for
coursework and exercises rather than for benchmarking or production systems.
The dataset deliberately contains one injected inconsistency; locating it is one of
the intended exercises (see The injected anomaly).
Overview
Property
Value
Examples… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/audio-3-speaker-dataset-v2.nepali-tts-synthetic-v2
Nepali TTS Synthetic v2
383,298 synthetic Nepali (ne) speech/text pairs, 24 kHz mono 16-bit WAV embedded
as-is (no re-encode, no resampling).
Generated by the synthetic_pipeline in milanakdj/TTS_training: Edge TTS
synthesis → optional voice conversion against a pool of 600 real multi-speaker
reference clips → ASR-based QC gate on character error rate.
Read this before training on it
Only 48% of rows are voice-converted. Each row carries a kept field
recording… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-tts-synthetic-v2.japanese-anime-speech-v2-splitdataset split from joujiboi/japanese-anime-speech-v2
twi_bible_v2_tts
Text-to-Speech Dataset
ewe_bible_v2_tts
Text-to-Speech
swiss-german-city-sentences_v2
Swiss German City Sentences v2
Synthetic Swiss German speech dataset with city name sentences across multiple dialects.
lfm2-tool-aware-dataset-v2
LFM2-Tool-Aware Dataset (v2)
Synthetic speech dataset for fine-tuning LFM2.5-Audio-class audio LLMs to handle both turns of a tool-augmented voice flow: acknowledge briefly on turn 1, then narrate the dispatcher's result on turn 2 after the coordinator injects it via set_context().
Used to train matbee/lfm2.5-audio-tool-aware-v2 (~97% accuracy on the eval split, including the new turn-2 narration class).
What's new in v2
The v1 dataset taught the model to ack-and-stop… See the full description on the dataset page: https://huggingface.co/datasets/matbee/lfm2-tool-aware-dataset-v2.GLaDOS-audio-v2
GLaDOS Audio v2 — Full Portal Wiki Collection
A complete collection of GLaDOS voice lines scraped from the
Portal Wiki,
covering all four canonical voice-line pages. Intended for TTS
fine-tuning and voice-cloning research.
Dataset Summary
Source
Description
portal1
Portal (2007) — all GLaDOS lines
portal2
Portal 2 (2011) — main campaign
portal2_coop
Portal 2 — Cooperative Testing Initiative
other
Other appearances (Poker Night at the Inventory… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/GLaDOS-audio-v2.bengali-telecom-customer-care-speech-v2
Bengali Telecom Customer Care Synthetic Speech Dataset v2
Dataset Description
This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts.
The dataset is intended for experiments with:
Bengali ASR/STT
Bengali TTS
Speech-to-text preprocessing
Telecom/customer-care domain adaptation
Synthetic speech research
This is a second version of the Bengali Telecom Customer Care Synthetic Speech Dataset. It follows the same… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech-v2.
