datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese-LiPS
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
⭐ Introduction
The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios.
🚀 Dataset Details
Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.common_voice_20_armenian
Common Voice 20 - Armenian
This dataset is the Armenian portion of Mozilla's Common Voice 20.0 release,
a massively multilingual collection of transcribed speech intended for speech technology research and development.
Dataset Details
Language: Armenian (hy)
Source: Mozilla Common Voice
Version: 20.0
License: CC0-1.0
ChildMandarin
ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
Introduction
ChildMandarin is a comprehensive, open-source Mandarin Chinese speech dataset specifically designed for research on young children aged 3 to 5. This dataset addresses the critical lack of publicly available resources for this age group, enabling advancements in automatic speech recognition (ASR), speaker verification (SV), and other related fields. The dataset is released… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/ChildMandarin.chinese-lips-speech-slide-probe
Chinese-LiPS Speech + Slide Probe
A self-contained probe set for testing whether visual slide context helps
simultaneous speech translation — with the input as audio, not transcripts.
Why audio matters: feeding a transcript to a text LLM deletes the acoustic
ambiguity (homophones, polysemy) that slide context is meant to resolve; the
transcript already commits to one reading. Any honest test of "does vision help
streaming ST" must consume speech.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.l2-arctic-manual-v5.0-16k
l2-arctic-manual-v5.0-16k
This dataset is a prepared derivative of L2-ARCTIC v5.0 that keeps only
the manually annotated material and converts the audio to 16 kHz mono FLAC.
It is designed to plug into the current peacock-asr training code, which
can consume a Hugging Face dataset with audio plus phonemes.
Included splits
train: 1800 rows, 1.84 hours
validation: 899 rows, 0.94 hours
test: 900 rows, 0.88 hours
suitcase: 22 rows, 0.44 hours
The scripted subset uses the… See the full description on the dataset page: https://huggingface.co/datasets/chikingsley/l2-arctic-manual-v5.0-16k.primewords_chinese_corpus_set_1chinese-lips-longform-debug
Chinese-LiPS Long-Form (zh long streaming speech)
Reconstructed continuous long-speech streams from
BAAI/Chinese-LiPS, for
slide-aware / streaming speech-translation development and evaluation. Each
source video (one speaker, one scripted lecture with slides) was released as
pre-segmented clips; here they are re-joined into the full talk.
Two variants of the same 3 talks (~97 min speech total):
config
how segments are placed
use
orig_timeline
at their original session… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-longform-debug.omniASR-igbo-blindspots
omniASR Igbo Blind Spot Dataset
Research Questions
This dataset investigates three interrelated questions about multilingual ASR performance on tonal languages:
Operational Definition: What does "language support" mean when a model lists 1,600+ languages? Does coverage imply functional accuracy on linguistically meaningful distinctions?
Diagnostic Validity: Can tonal diacritic preservation serve as a diagnostic for acoustic competence vs. orthographic pattern matching… See the full description on the dataset page: https://huggingface.co/datasets/Chiz/omniASR-igbo-blindspots.chichewa-trigrams-speech-text-parallel
Chichewa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 132549 parallel speech-text pairs for Chichewa, a language spoken primarily in Malawi. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Chichewa - ny
Task: Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-trigrams-speech-text-parallel.Chinese-Speech-Dataset
🎧 Chinese (Simplified) Speech Dataset
The Chinese (Simplified) speech dataset is a high-quality speech audio dataset developed to support scalable AI and machine learning solutions with diverse and structured audio data. It contains 105 hours of speech data across 700 audio files, delivered in MP3 and WAV formats, with a total size of 229 MB. This well-balanced audio dataset provides reliable voice data, featuring 54% female and 46% male speakers, with age distribution ranging from… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Chinese-Speech-Dataset.free_st_chinese_mandarin_corpusChildren_Counsel
아동·청소년 상담 데이터셋 (Children Counseling Dataset)
This dataset contains counseling data for children and adolescents, including both audio recordings and transcriptions.
Dataset Structure
The dataset is organized as follows:
audio/: Contains the audio recordings of counseling sessions in MP3 format
data/: Contains JSON files with transcriptions and metadata for each session
Usage
This dataset can be used for:
Training speech recognition models for counseling… See the full description on the dataset page: https://huggingface.co/datasets/ironDong/Children_Counsel.omnilora-kazakh-child-mvp
OmniLoRA Kazakh Child-Voice TTS — Cleaned & Emotion-Labeled Subset
A 535-clip Kazakh child-speech subset derived from
galammadin-asr/child-asr-kazakh,
cleaned through a 3-stage automatic filter and hand-labeled with one of six
emotion categories. Built for fine-tuning a LoRA adapter on top of
OmniVoice (Method 5 of a 5-method Kazakh
TTS benchmark, CSCI 595 final project).
Dataset summary
Clips
535
Language
Kazakh (kk)
Sample rate
16 kHz (inherits from… See the full description on the dataset page: https://huggingface.co/datasets/neversi123/omnilora-kazakh-child-mvp.chichewa_english_code_switch_dataset
Chichewa-English Code-Switched Speech Dataset
Dataset Description
A speech dataset containing 247 audio recordings of Chichewa-English code-switched phrases. Code-switching — the practice of alternating between two or more languages within a single conversation — is extremely common in Malawi and across multilingual African communities. This dataset captures that natural linguistic behavior in spoken form.
Purpose
This dataset is designed to support research… See the full description on the dataset page: https://huggingface.co/datasets/suru8-ai/chichewa_english_code_switch_dataset.Chichewa-Synthetic-ASR-DatasetSynthetic Chichewa ASR dataset generated using a fine-tuned version of the YourTTS model.
Sample rate: 24kHz.
Total duration: 550 hours.
Chichewa-Speech-Dataset
🎧 Chichewa Speech Dataset
The Chichewa Speech Dataset is a high-quality speech audio dataset designed to provide structured and scalable audio data for AI-driven voice technologies. It includes 94 hours of audio data across 740 files, delivered in MP3 and WAV formats, with a total size of 272 MB. This well-organized audio dataset ensures balanced and diverse voice data, with 52% female and 48% male speakers, and a wide age distribution from 18 to 50+ years. The dataset language is… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Chichewa-Speech-Dataset.YodaLingua-Chinese
YodaLingua-Chinese
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Chinese portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
86,229 audio–transcription pairs
Total duration
235.5 hours
Speakers
1,873 distinct speakers
Audio format
MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Chinese.CHILDES-Aligned
[!IMPORTANT]
How to access this dataset: the official public release is hosted by TalkBank at
https://talkbank.org/childes/access/Derived/CHILDES-Aligned.html (audio archives +
CSV/JSONL metadata, CC BY-NC-SA 4.0). Please obtain the dataset there.
This Hugging Face copy is retained gated, for internal use; access requests are
approved manually and general requests may be declined — use the TalkBank release instead.
CHILDES-Aligned: Curated Child-Speech Dataset (BEACON)
English… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/CHILDES-Aligned.CoReBench_v1
Dataset Card for CoReBench_v1
Dataset Summary
COREBench is a comprehensive conversational reasoning benchmark designed to evaluate audio language models on reasoning capabilities in multi-turn conversations.
Example Instance
Question: What is the fruit the first speaker likes most?
Audio Sample:
Download audio: https://huggingface.co/datasets/chiheemwong/CoReBench_v1/audio/ebd9de53fbca567cf675.mp3
[Transcript]
Zinaida: Alright team, let's nail this chorus. We… See the full description on the dataset page: https://huggingface.co/datasets/chiheemwong/CoReBench_v1.
