CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01disco-eth /WorldSpeech WorldSpeech A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR estimate, and four DNSMOS-P.835 quality scores. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/WorldSpeech.audioautomatic-speech-recognition10M<n<100M49 likes45k downloads4mo agoHugging Face02disco-eth /EuroSpeech EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. This dataset is 16 kHz, the 24 kHz version of EuroSpeech can be found at… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech.audioautomatic-speech-recognition10M<n<100M99 likes36k downloads5mo agoHugging Face03disco-eth /EuroSpeech-24kHz EuroSpeech 24 kHz Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. Dataset Summary Languages: 22 European languages (see detailed… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech-24kHz.audioautomatic-speech-recognition10M<n<100M3 likes1.7k downloads5mo agoHugging Face04Benji-fish /ethiopian-languages-speech-dataset Leyu Ethiopian Languages Speech Dataset Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya. Dataset Summary Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti) Total examples: 2750 License: CC-BY-4.0 Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.audioautomatic-speech-recognition1K<n<10K0 likes982 downloads1mo agoHugging Face05badrex /ethiopian-speech-flat Ethio Speech Copus — Afrivoices Ethiopian 📌 Overview The Ethio Speech Corpus dataset is a multilingual speech corpus containing audio–text pairs across five Ethiopian languages. It is designed to support the development of speech-to-text technologies for low-resource languages. This dataset is part of the Afrivoices initiative — a collaborative effort to create a large-scale ASR dataset for African languages. The broader goal of the initiative is to collect 600 hours… See the full description on the dataset page: https://huggingface.co/datasets/badrex/ethiopian-speech-flat.audioautomatic-speech-recognition100K<n<1M2 likes279 downloads8mo agoHugging Face06hadamard-2 /fleurs-ethiopian-v2 FLEURS — Ethiopian Languages This dataset is a restructured v2 conversion of the Google FLEURS dataset for two Ethiopian languages: Amharic (am_et) and Oromo (om_et). Subsets Subset Language ISO 639-2 Train Dev Test amh Amharic amh 3,163 223 516 orm Oromo orm 1,701 19 41 Splits Split Description train Training split dev Development split (renamed from validation in original FLEURS) test Test split Usage from… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/fleurs-ethiopian-v2.audioautomatic-speech-recognition1K<n<10K0 likes172 downloads7mo agoHugging Face07LeyuCompetition /benji-ethiopian-languages-speech-dataset Leyu Ethiopian Languages Speech Dataset Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya. Dataset Summary Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti) Total examples: 2750 License: CC-BY-4.0 Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/benji-ethiopian-languages-speech-dataset.audioautomatic-speech-recognition1K<n<10K0 likes71 downloads21d agoHugging Face08Ethan615 /taiwan-conversation-context-100-domainsgated Taiwan Conversation Context 100 Domains Dataset Description Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。 本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。 資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於: 語音生成資料前處理 Text-to-Speech, TTS Spoken Dialogue Generation Conversational AI Customer Service Dialogue Modeling Role-play Dialogue Dataset 台灣繁體中文語音模型訓練 生活情境問答模型訓練 對話式 AI 助理訓練 RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.texttext-generation1M<n<10M2 likes37 downloads5mo agoHugging Face09hadamard-2 /common-voice-24-ethiopian-v2 Common Voice 24.0 — Ethiopian Languages (v2) This is a cleaned and restructured version of hadamard-2/common-voice-24-ethiopian, which is a verbatim archival upload of the Mozilla Common Voice 24.0 scripted speech data for Amharic and Tigrinya. This version conforms to the Waxal ASR schema. Subsets Subset Language ISO 639-2 Clips Validated Hours Speakers amh Amharic amh 1,045 1.82h 46 tir Tigrinya tir 69 0.10h 16 Splits Split… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/common-voice-24-ethiopian-v2.audioautomatic-speech-recognition1K<n<10K0 likes17 downloads7mo agoHugging Face10hadamard-2 /common-voice-24-ethiopian Common Voice 24.0 — Ethiopian Languages This dataset is a verbatim archival upload of the Mozilla Common Voice 24.0 scripted speech data for two Ethiopian languages: Amharic (am) and Tigrinya (ti), sourced from the Mozilla Data Collective. Subsets Subset Language Code Clips Total Hours Validated Hours Speakers amharic Amharic am 1,632 2.85h 1.82h 46 tigrinya Tigrinya ti 451 0.65h 0.10h 16 Splits Each subset contains the following splits:… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/common-voice-24-ethiopian.audioautomatic-speech-recognition1K<n<10K0 likes13 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.