datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese-LiPS
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
⭐ Introduction
The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios.
🚀 Dataset Details
Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.lex_fridman_podcast
Dataset Card for "lex_fridman_podcast"
Dataset Summary
This dataset contains transcripts from the Lex Fridman podcast (Episodes 1 to 325).
The transcripts were generated using OpenAI Whisper (large model) and made publicly available at: https://karpathy.ai/lexicap/index.html.
Languages
English
Dataset Structure
The dataset contains around 803K entries, consisting of audio transcripts generated from episodes 1 to 325 of the Lex Fridman… See the full description on the dataset page: https://huggingface.co/datasets/nmac/lex_fridman_podcast.Afar-language-text-to-speech-TTS
Usage
This dataset is designed to support the development of Text-to-Speech (TTS) systems for the Afar language. It can be integrated into web applications, mobile apps, desktop software, or other platforms that require natural-sounding Afar voice synthesis or accurate spoken language recognition.
For applications involving virtual avatars or voice personas, the following culturally appropriate voice names are recommended:
Female Voices: Emeli, Hanaawi, Kareera, Laysani, Kulsuma… See the full description on the dataset page: https://huggingface.co/datasets/Charif-Ayfarah/Afar-language-text-to-speech-TTS.lost-in-speech
Lost in Speech
A trilingual benchmark for reference-free classification of synthetically introduced factual and contextual alterations in English, Russian, and Kazakh. It contains 12,013 samples derived from news articles, with text, synthesized speech, and ASR transcript representations used in the study.
Altered samples are LLM-generated rewrites with a controlled alteration type—contradiction, fabrication, or context inconsistency—and severity level—mild, moderate, or severe.… See the full description on the dataset page: https://huggingface.co/datasets/maristombayeva/lost-in-speech.2026-dwesui-g01-neurologia
DWESUI 2026 - Grupa 1 - neurologia
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 1 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Domena: neurologia
Licencja zrodla: nagrania YouTube CC-BY + synteza TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował zbiór… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g01-neurologia.Latvian-Speech-Dataset
Latvian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Latvian (la)
🏷️ Tags
Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Latvian
📦 Size Category
n < 1K
2026-dwesui-g02-kulinarna
DWESUI 2026 - Grupa 2 - kulinarna (PIEROGA)
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 2 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna
Domena: kulinarna
Licencja zrodla: nagrania YouTube CC-BY/CC-BY-SA + TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g02-kulinarna.last
last
This is a merged speech dataset containing 345 audio segments from 2 source datasets.
Dataset Information
Total Segments: 345
Speakers: 7
Languages: en
Emotions: neutral, angry, happy, sad
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/last.Lithuanian-Speech-Dataset
Lithuanian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Lithuanian (lt)
🏷️ Tags
Lithuanian, Audio, Speech, Speech Recognition, ML, Machine, Machine Learning
📦 Size Category
n < 1K
Luxembourgish-Speech-Dataset
Luxembourgish Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Luxembourgish (lb)
🏷️ Tags
Audio, Speech, Speech Recognition, Machine, Machine Learning, ML
📦 Size Category
n < 1K
WildVid-LIP
WildVid-LIP: In-The-Wild Temporal Anchors for Visual Speech Recognition
WildVid-LIP is a large-scale, open-source dataset mapping over 100,000 curated temporal segments from unconstrained, real-world YouTube videos. It provides precise timestamp anchors optimized for training Visual Speech Recognition (VSR / Lip-Reading), audio-visual synchronization, and multimodal self-supervised models.
Instead of distributing heavy, monolithic video files—which introduces platform friction… See the full description on the dataset page: https://huggingface.co/datasets/Rizul2159/WildVid-LIP.luel-multilingual-tts-samples
Multilingual TTS Samples (Luel)
License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE.
A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.Lingala-Speech-Dataset
Lingala Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Lingala (ln)
🏷️ Tags
Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Lingala
📦 Size Category
n < 1K
LibriReplay-DOA
LibriReplay-DOA (Anonymous Submission)
Overview
LibriReplay-DOA is a multi-channel multi-speaker replay dataset designed for evaluating
robust speech processing systems under realistic room playback conditions.
The dataset contains replay recordings captured in real rooms under
multiple playback configurations (DOA settings). Each session includes
multiple overlapping speakers.
This dataset is released for peer-review purposes.
Dataset Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/real-recordings/LibriReplay-DOA.
