datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxlingua107_wds
VoxLingua107
VoxLingua107 is a speech dataset for training spoken language identification models.
The dataset consists of short speech segments automatically extracted from YouTube videos and labeled according the language of the video title and description, with some post-processing steps to filter out false positives.
VoxLingua107 contains data for 107 languages. The total amount of speech in the training set is 6628 hours.
The average amount of data per language is 62 hours.… See the full description on the dataset page: https://huggingface.co/datasets/TalTechNLP/voxlingua107_wds.mls_hq_urgent_track1behavior-sd
🎙️ Behavior-SD
Official repository for our NAACL 2025 paper:Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language ModelsSehun Lee*, Kang-wook Kim*, Gunhee Kim (* Equal contribution)
🏆 SAC Award Winner in Speech Processing and Spoken Language Understanding
🔗 Links
Project Page
Code
📖 Overview
We explores how to generate natural, behaviorally-rich full-duplex spoken dialogues using large language models (LLMs).
We introduce:… See the full description on the dataset page: https://huggingface.co/datasets/yhytoto12/behavior-sd.CounterStrike-1K-360-wds
CounterStrike-1K — 360p WebDataset shards
This repo contains the 360p shards of CounterStrike-1K. Use the main repo to browse the manifest, schema, and subsets.
360p is the recommended resolution for most training pipelines — the actions/state/events/metadata sidecars are identical to the 720p shards, so you can swap resolutions without touching downstream code.
Quickstart
Start a fresh uv project and add the loader:
mkdir cs1k-demo && cd cs1k-demo
uv init
uv add… See the full description on the dataset page: https://huggingface.co/datasets/ArnieRamesh/CounterStrike-1K-360-wds.Galgame_Speech_SER_16kHz
Dataset Card for Galgame_Speech_SER_16kHz
[!IMPORTANT]The following rules (in the original repository) must be followed:
必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关!
训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。
English:
You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_SER_16kHz.audiosnippets_long_1Munsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1multi_round_speech_180kGalgame_Speech_ASR_16kHz
Dataset Card for Galgame_Speech_ASR_16kHz
[!IMPORTANT]The following rules (in the original repository) must be followed:
必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关!
训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。
English:
You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_ASR_16kHz.log_prompt_datadns5-16k
DNS5 16kHz
Resampled subset of the ICASSP 2022 DNS Challenge dataset.
All audio files resampled from 48kHz to 16kHz and stored as FLAC (lossless compression),
packed into tar shards.
Structure
clean/shard_0000.tar # Clean speech (VCTK and other corpora)
clean/shard_0001.tar
...
noise/shard_0000.tar # Environmental noise (AudioSet, Freesound)
...
impulse_responses/shard_0000.tar # Room impulse responses
...
Each tar contains FLAC files with their… See the full description on the dataset page: https://huggingface.co/datasets/richiejp/dns5-16k.laion-coco-13m-tarmajestrino-1.00-16xk5-sae-features
Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5)
Top-2000 activating audio samples for each feature in the
Majestrino 1.00 SAE.
Overview
Metric
Value
SAE Architecture
16x expansion, k=5, d_model=768
Total Features
12,288
Alive Features
10,684
Audio per Feature
Up to 2,000 highest-activating
Audio Format
Opus (24 kbps OGG container)
Total TAR Files
1069
Source Dataset
laion/majestrino-data
File Structure
Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.SpokenWOZ-Train-Audiovoa_myanmar_asr_audio_1
📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology.
Overview
This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.emo_speech_filtered_v12 second filtered emotional speech in webdataset format
https://huggingface.co/datasets/EQ4You/Emotional_Speech
LJSpeech-1.1
The LJ Speech Dataset
Version 1.1
July 5, 2017
https://keithito.com/LJ-Speech-Dataset
OVERVIEW
This is a public domain speech dataset consisting of 13,100 short audio clips
of a single speaker reading passages from 7 non-fiction books. A transcription
is provided for each clip. Clips vary in length from 1 to 10 seconds and have
a total length of approximately 24 hours.
The texts were published between 1884 and 1964, and are in the public domain.
The audio was recorded in… See the full description on the dataset page: https://huggingface.co/datasets/badayvedat/LJSpeech-1.1.AudioQA-1M115hours_pvtv_myanmar_asr
115 Hours PVTV Myanmar ASR
This dataset contains 156,262 audio-transcript pairs of spoken Burmese, totaling approximately 115.31 hours. The audio segments were extracted from publicly available YouTube videos published by PVTV and aligned using subtitle timestamps.
Dedication
This dataset would not exist without the persistent voices of PVTV editors, journalists, narrators, and production teams, who continue to speak to the people under difficult conditions. PVTV is the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/115hours_pvtv_myanmar_asr.vocal_bursts_taxonomy_100_clean_wdsru-book-mix-10h
ru-book-mix-10h
A 10-hour synthetic Russian-audiobook diarization benchmark. 600 one-minute
FLAC clips (16 kHz mono, 16-bit, lossless) with NIST RTTM ground truth, generated by
mexus/diarization-benchmark
from its5Q/biggest-ru-book (speech) and bilguun/musan-noise (background).
Intended use: diarization evaluation only. This dataset is not
suitable for training — the same source voices repeat across files, so any
model that trains on it will leak voice identity into its test split.… See the full description on the dataset page: https://huggingface.co/datasets/mexus/ru-book-mix-10h.stortinget_speech_corpus_v1.0
Dataset Card for Stortinget Speech Corpus V1.0
Overview
This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability.
The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.AVQA-R1-6KThis repository contains data presented in EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning.
For training and inference, please refer to the Code: https://github.com/HarryHsing/EchoInk
Data Format in AVQA-R1-6K:
{
"problem_id": 0,
"problem": "What is the source of the sound in the video?",
"data_type": "image_audio",
"problem_type": "multiple choice",
"options": [
"A. motorcycle",
"B. automobile"… See the full description on the dataset page: https://huggingface.co/datasets/harryhsing/AVQA-R1-6K.common_voice_16_1_fr_smallEmilia-YODAS-KO-filteredlibritts-r-gzVoxceleb1voxceleb2-40k-part1-preprocess-all-files-separateurdu-turn-detection-audio-v2
🗣️ Urdu Turn Detection (Audio Dataset V2)
This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech.
It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection.
🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-turn-detection-audio-v2.Japanese-Eroge-Voice
Japanese-Eroge-Voice
Description
This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model.
Preprocessing Steps
The raw audio data has undergone the following preprocessing steps:
Loudness Normalization:
Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/test12313/Japanese-Eroge-Voice.
