datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepDialogue-orpheus
DeepDialogue-orpheus
DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text.
🚨 Important Notice
This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.or_in_datasetYouTube-Cantonese
Cantonese Audio Dataset from YouTube
This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models.
Data Source and Processing
The data was obtained through the following process:
Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.latam-spanish-speech-orpheus-tts-24khz
LATAM Spanish High-Quality Speech Dataset (24kHz - Orpheus TTS Ready)
Dataset Description
This dataset contains approximately 24 hours of high-quality speech audio in Latin American Spanish, specifically prepared for Text-to-Speech (TTS) applications like OrpheusTTS, which require a 24kHz sampling rate.
The audio files are derived from the Crowdsourced high-quality speech datasets made by Google and were obtained via OpenSLR. The original recordings were high-quality… See the full description on the dataset page: https://huggingface.co/datasets/GianDiego/latam-spanish-speech-orpheus-tts-24khz.afaan-oromoo-speech
Dataset.ET Afaan Oromoo Speech — v0.1.0
9.843 hours · 3,594 clips · 74 speakers · 3,283 distinct prompts
Dataset Summary
Read speech in Afaan Oromoo, crowdsourced from volunteer contributors in Ethiopia
through a Telegram bot, peer-validated by other contributors, and screened
acoustically before release. Afaan Oromoo has very little open speech data; this
corpus exists to change that.
Contributors read a displayed prompt aloud, other contributors listen and vote… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/afaan-oromoo-speech.Thai-Food-Ordering-Dataset
🍲 Thai Food Ordering Speech Dataset
A Specialized Speech Recognition Corpus for Thai Food Ordering and Restaurant Contexts
📌 Dataset Overview
The Thai Food Ordering Speech Dataset is a domain-specific audio dataset created to develop and enhance Automatic Speech Recognition (ASR) systems, specifically targeting Thai food ordering in food courts, street stalls, and dining environments.
In real-world food court operations, manual order… See the full description on the dataset page: https://huggingface.co/datasets/KittipatPaisanpudinun/Thai-Food-Ordering-Dataset.Magpie-Speech-Orpheus-125k
Magpie-Speech-Orpheus-125k
A ~125k-sample synthetic speech dataset generated by applying the Magpie instruction-synthesis approach to the Orpheus-TTS LLM-based text-to-speech model, then decoding audio tokens with the SNAC 24 kHz codec.
Blog (EN): https://huggingface.co/blog/Aratako/magpie-speech
Blog (JA): https://zenn.dev/aratako_lm/articles/87d8988d44ba4d
This dataset is entirely synthetic: text prompts and audio tokens were produced by Orpheus-TTS and decoded to waveforms via… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Magpie-Speech-Orpheus-125k.original-songs
Dataset Card for "original-songs" (Audio + análisis DSP)
Dataset Summary
Dataset pequeño de canciones originales creadas con IA, cada una con su WAV,
letra transcrita automáticamente (Whisper) y un análisis DSP completo (tempo,
tonalidad, loudness, features perceptuales) además de detección de contenido
explícito. Pensado para quien quiera mejorar modelos open source: extracción
de features musicales, clasificación de audio, transcripción y moderación de
letras.… See the full description on the dataset page: https://huggingface.co/datasets/arnauquest/original-songs.NPSC_ortoThe Norwegian Parliament Speech Corpus (NPSC) is a corpus for training a Norwegian ASR (Automatic Speech Recognition) models. The corpus is created by Språkbanken at the National Library in Norway.
NPSC is based on sound recording from meeting in the Norwegian Parliament. These talks are orthographically transcribed to either Norwegian Bokmål or Norwegian Nynorsk. In addition to the data actually included in this dataset, there is a significant amount of metadata that is included in the original corpus. Through the speaker id there is additional information about the speaker, like gender, age, and place of birth (ie dialect). Through the proceedings id the corpus can be linked to the official proceedings from the meetings.
The corpus is in total sound recordings from 40 entire days of meetings. This amounts to 140 hours of speech, 65,000 sentences or 1.2 million words.
This dataset builds on this corpus. In addition it adds two columns with machine generated orthographic text.oral-arguments-us
US Court Oral Arguments -- metadata and transcripts
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Source & credit — Free Law Project / CourtListener
Every recording catalogued here was collected and catalogued by CourtListener, and this dataset is
sliced… See the full description on the dataset page: https://huggingface.co/datasets/docketx/oral-arguments-us.ivan_shamyakin_tryvozhnae_shchastse_output_original
Трывожнае шчасце — арыгінальнае аўдыё
Аўтар / Author: Іван ШамякінМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
ivan_shamyakin_tryvozhnae_shchastse_output
Доўгасць аўдыё
30h45m
Радкоў у датасеце
6,523
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ivan_shamyakin_tryvozhnae_shchastse_output_original.kuzma_chorny_zyamlya_output_original
Зямля — арыгінальнае аўдыё
Аўтар / Author: Кузьма ЧорныМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
kuzma_chorny_zyamlya_output
Доўгасць аўдыё
28h15m
Радкоў у датасеце
6,820
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/kuzma_chorny_zyamlya_output_original.sravaani-indic-diarbench-oracle-v1
SraVaani Indic DiarBench Oracle ASR
This is a portable evaluation-only oracle-turn view derived from
sarvamai/indic-diarbench at the
immutable revision 92877bad8aab6e598167d91c6ee02aa8ca6ede09. It contains Hindi and Telugu only.
Do not use these turns for fine-tuning if Indic DiarBench will remain an
external benchmark. Training on this export contaminates the test set.
Configurations
Config
Test rows
Audio hours
Purpose
primary
2,417
3.8233
Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhirl/sravaani-indic-diarbench-oracle-v1.reamixed-original-tracks
reamixed_original_tracks
Extracted and organized original tracks for ReaMIXed contests.
Summary
Included files: 6572
Included size: 72.09 GB
Known missing items
202301: official reamixed_original_tracks link is no longer valid; local original tracks archive is missing.
202603: present in latest metadata but no local archive/extracted folder is available yet.
YouTube-English
English Audio Dataset from YouTube
This dataset contains English audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models.
Data Source and Processing
The data was obtained through the following process:
Download: Audio (.m4a) and available English subtitles (.srt for en, en.j3PyPqV-e1s) were… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-English.kuzma_chorny_poshuki_buduchyni_output_original
Пошукі будучыні — арыгінальнае аўдыё
Аўтар / Author: Кузьма ЧорныМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
kuzma_chorny_poshuki_buduchyni_output
Доўгасць аўдыё
22h58m
Радкоў у датасеце
5,530
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/kuzma_chorny_poshuki_buduchyni_output_original.uladzimir-karatkevich-dzikae-paliavanne-karalia-stakha-aleg-garbuz-output_original
Дзікае паляванне Кароля Стаха — арыгінальнае аўдыё
Аўтар / Author: Уладзімір КараткевічМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
uladzimir-karatkevich-dzikae-paliavanne-karalia-stakha-aleg-garbuz-output
Доўгасць аўдыё
7h37m
Радкоў у датасеце
2,585
Структура… See the full description on the dataset page: https://huggingface.co/datasets/fosters/uladzimir-karatkevich-dzikae-paliavanne-karalia-stakha-aleg-garbuz-output_original.yakub_kolas_novaya_zyamlya_output_original
Новая зямля — арыгінальнае аўдыё
Аўтар / Author: Якуб КоласМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
yakub_kolas_novaya_zyamlya_output
Доўгасць аўдыё
9h49m
Радкоў у датасеце
1,434
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/yakub_kolas_novaya_zyamlya_output_original.wayuu_CO_test
Dataset Audio Duration
The dataset consists of 810 audio recordings, each accompanied by its respective transcription. The lexical corpus encompasses approximately 1,000 unique words.
Total Audio Duration: 2801 seconds (approximately 34 minutes)
Average Audio Duration: 3.41 seconds
The dataset offers valuable insights into the Wayuunaiki language's phonetic and linguistic characteristics. It's important to note that the dataset originates from recordings and transcriptions of the… See the full description on the dataset page: https://huggingface.co/datasets/orkidea/wayuu_CO_test.Soblogeryh_raspe_prygody_barona_myunhau_zena_output_original
Прыгоды барона Мюнхаўзена — арыгінальнае аўдыё
Аўтар / Author: Эрых РаспэМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
eryh_raspe_prygody_barona_myunhau_zena_output
Доўгасць аўдыё
1h58m
Радкоў у датасеце
494
Структура
Кожны радок змяшчае:
audio — арыгінальны… See the full description on the dataset page: https://huggingface.co/datasets/fosters/eryh_raspe_prygody_barona_myunhau_zena_output_original.ales_zhuk_praklytaya_lyubow_output_original
Пракляты любоў — арыгінальнае аўдыё
Аўтар / Author: Алесь ЖукМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
ales_zhuk_praklytaya_lyubow_output
Доўгасць аўдыё
3h33m
Радкоў у датасеце
854
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ales_zhuk_praklytaya_lyubow_output_original.vasil_bykau_output_original
Зборнік — арыгінальнае аўдыё
Аўтар / Author: Васіль БыкаўМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
vasil_bykau_output
Доўгасць аўдыё
6h02m
Радкоў у датасеце
1,402
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя
chunk_uid —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/vasil_bykau_output_original.leyu-oromo-speech-corpus-2026
Leyu Afaan Oromo Speech Corpus 2026
Official speech dataset submission for the Leyu Data Collection Competition 2026.
Organization & Team
Hugging Face Org: SoundWaveET
Dataset Repo: SoundWaveET/leyu-oromo-speech-corpus-2026
yanka_sipakou_odzium_output_original
Адзіум — арыгінальнае аўдыё
Аўтар / Author: Янка СіпакоўМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
yanka_sipakou_odzium_output
Доўгасць аўдыё
1h07m
Радкоў у датасеце
273
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя
chunk_uid… See the full description on the dataset page: https://huggingface.co/datasets/fosters/yanka_sipakou_odzium_output_original.iakub-kolas-kazki-zhytstsia-output_original
Казкі жыцця — арыгінальнае аўдыё
Аўтар / Author: Якуб КоласМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
iakub-kolas-kazki-zhytstsia-output
Доўгасць аўдыё
1h49m
Радкоў у датасеце
504
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/iakub-kolas-kazki-zhytstsia-output_original.viktar_prau_dzin_output_original
Зборнік — арыгінальнае аўдыё
Аўтар / Author: Віктар ПраўдзінМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
viktar_prau_dzin_output
Доўгасць аўдыё
4h10m
Радкоў у датасеце
1,008
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/viktar_prau_dzin_output_original.nepal-oral-packs-open
Nepal Oral Open Packs
Public educational packs for the Nepal Oral Train app and NGO classroom use (AkAiNp).
Versioned lesson packs (meanings, tags, provisional text)
Seed / model-speaker audio only when license is open
Language origin metadata (family, homeland, script status, varieties)
Multi-track over time; starts empty until steward-approved open content exists
This is what the mobile app may download read-only after a pack is published.It is not a dump of every phone… See the full description on the dataset page: https://huggingface.co/datasets/AkAiNp/nepal-oral-packs-open.genadz_pashkou_output_original
Зборнік — арыгінальнае аўдыё
Аўтар / Author: Генадзь ПашкоўМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
genadz_pashkou_output
Доўгасць аўдыё
0h28m
Радкоў у датасеце
101
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя
chunk_uid —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/genadz_pashkou_output_original.yanka_bryl_ptushki_i_gne_zdy_output_original
Птушкі і гнёзды — арыгінальнае аўдыё
Аўтар / Author: Янка БрыльМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
yanka_bryl_ptushki_i_gne_zdy_output
Доўгасць аўдыё
18h52m
Радкоў у датасеце
4,027
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/yanka_bryl_ptushki_i_gne_zdy_output_original.
