CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lingamvamshikrishnareddy /ramanv-tts-all-rawgated ramanv-tts-all-raw Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages. textautomatic-speech-recognition1M<n<10M0 likes4.7k downloads9d agoHugging Face02TheNHz /ellipsis-lrs3-rawgated LRS3-TED — verified mirror A mirror of the LRS3-TED dataset (Lip Reading Sentences 3), preserved because the official distribution has been discontinued. This repository adds no new data: it is a re-hosted copy with a full verification report against the official file list, so you know exactly what is and is not here. Attribution LRS3-TED was created by Triantafyllos Afouras, Joon Son Chung and Andrew Zisserman (Visual Geometry Group, University of Oxford): T.… See the full description on the dataset page: https://huggingface.co/datasets/TheNHz/ellipsis-lrs3-raw.textautomatic-speech-recognition1K<n<10K8 likes786 downloads2mo agoHugging Face03aranemini /northern-kurdish-raw-audio Northern Kurdish Raw Audio Collection Overview This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources. The collection was assembled to support research and development in: Automatic Speech Recognition (ASR) Speech Translation (ST) Text-to-Speech (TTS) Self-supervised Learning (SSL) Spoken Language Understanding (SLU) The dataset contains more than 2,000 hours… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-raw-audio.audioautomatic-speech-recognition1K<n<10K0 likes759 downloads3mo agoHugging Face04serdarcaglar /turkish-audiobook-rawgated Turkish Audiobook Speech Corpus (Raw) Türkçe konuşma araştırmaları için derlenmiş, işlenmemiş uzun-form ses kayıtlarından oluşan bir koleksiyon. Kayıtlar çeşitli kaynaklardan bir araya getirilmiştir ve konuşmacı, kayıt ortamı, süre ve ses kalitesi bakımından geniş bir çeşitlilik gösterir. İçerik Uzun-form Türkçe konuşma kayıtları (m4a / mp3) Kaynağa göre klasörlenmiş düz dizin yapısı Transkript, hizalama veya segmentasyon içermez — ham hâldedir… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-audiobook-raw.audiotext-to-speechn<1K0 likes388 downloads25d agoHugging Face05aranemini /southern-kurdish-raw-audio Southern Kurdish Raw Audio Collection Overview This repository contains approximately 170 hours of Southern Kurdish (SDH) raw speech collected from publicly available media sources, podcasts, interviews, news broadcasts, and online programs. The main sources are Aryen TV and Kurd Channel. The recordings mainly consist of spontaneous and semi-spontaneous speech, covering diverse speakers, topics, and acoustic conditions. While the majority of the content is… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/southern-kurdish-raw-audio.audioautomatic-speech-recognitionn<1K0 likes190 downloads3mo agoHugging Face06aranemini /central-kurdish-audiobook-raw Central Kurdish Audiobook Raw Audio Collection Overview This repository contains a large collection of raw Central Kurdish (Sorani Kurdish) audiobook recordings gathered from publicly available online sources. The collection was assembled to support research and development in: Automatic Speech Recognition (ASR) Speech Translation (ST) Text-to-Speech (TTS) Self-supervised learning The dataset contains approximately 4,300 hours of speech collected from 1026… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-audiobook-raw.audioautomatic-speech-recognition10K<n<100K1 likes172 downloads3mo agoHugging Face07Menlo /raw-speech-whispervq-v1 Dataset Overview This dataset contains over 2,4M English ASR samples, using: The a training set of parler-tts/mls_eng_10k Tokenized using WhisperVQ. Usage from datasets import load_dataset, Audio # Load Instruction Speech dataset dataset = load_dataset("homebrewltd/raw-speech-whispervq-v1",split='train') Dataset Fields Field Type Description tokens sequence Tokenized using Encodec text sequence Converted audio tokens Bias, Risks… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/raw-speech-whispervq-v1.textautomatic-speech-recognition1M<n<10M0 likes157 downloads2y agoHugging Face08oddadmix /arabic-audio-collection-algerian-rawi Rawi Postcast Arabic Speech Dataset Dataset Summary The Rawi Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 51 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-rawi.audiotext-to-speech1K<n<10K0 likes143 downloads3mo agoHugging Face09aranemini /hawrami-kurdish-raw-audio Hawrami Raw Audio Collection Overview This repository contains approximately 500 hours of Hawrami Kurdish raw speech collected from publicly available media sources. The dataset was gathered primarily from the Rocyar program broadcast on Sterk TV, along with additional publicly available Hawrami-language content. The recordings mainly consist of spontaneous and semi-spontaneous speech, including interviews, discussions, cultural programs, storytelling, and other… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/hawrami-kurdish-raw-audio.audioautomatic-speech-recognitionn<1K0 likes127 downloads3mo agoHugging Face10kiranpantha /no-filter-raw-NepaliParliamentDSv2audioautomatic-speech-recognition10K<n<100K0 likes58 downloads1y agoHugging Face11freococo /raw_1hr_myanmar_asr_audio 🇲🇲 Raw 1-Hour Burmese ASR Audio Dataset A 1-hour dataset of Burmese (Myanmar language) spoken audio clips with transcripts, curated from official public-service media broadcasts by PVTV Myanmar — the media voice of Myanmar’s National Unity Government (NUG). This dataset is intended for automatic speech recognition (ASR) and Burmese speech-processing research. ➡️ Author: freococo➡️ License: MIT➡️ Language: Burmese (my) 📦 Dataset Summary Duration: ~1 hour Chunks:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/raw_1hr_myanmar_asr_audio.audioautomatic-speech-recognitionn<1K0 likes41 downloads1y agoHugging Face12Quran-Lab /quranic-asr-cloud-rawdatagated Quranic ASR Provider Benchmark Results Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark. This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio. What Is Included Area Path Purpose Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.tabularautomatic-speech-recognition1K<n<10K1 likes32 downloads27d agoHugging Face13freococo /2hr_myanmar_asr_raw_audio 🇲🇲 Raw 2-Hour Burmese ASR Audio Dataset A ~2-hour Burmese (Myanmar language) ASR dataset featuring 1,612 audio clips with aligned transcripts, curated from official public-service educational broadcasts by FOEIM Academy — a civic media arm of FOEIM.ORG, operating under the Myanmar National Unity Government (NUG). This dataset is MIT-licensed as a public good — a shared asset for the Burmese-speaking world. It serves speech technology, education, and cultural preservation efforts… See the full description on the dataset page: https://huggingface.co/datasets/freococo/2hr_myanmar_asr_raw_audio.audioautomatic-speech-recognition1K<n<10K0 likes24 downloads1y agoHugging Face14Anamavajra-Labs /tantraloka-dyczkowski-rawgated Two views of the same 5,146 verses structured is the full record — 24 columns over all 37 āhnikas, including extracted entities, cross-references, parallel passages and audio alignment. Several of those columns hold JSON documents running to thousands of characters, which is what makes the dataset viewer unreadable in a browser: a row is a wall of text. reading (the default config) is a projection of the same rows onto the eight columns a reader wants — volume, chapter, verse… See the full description on the dataset page: https://huggingface.co/datasets/Anamavajra-Labs/tantraloka-dyczkowski-raw.audiotext-generation10K<n<100K0 likes21 downloads1mo agoHugging Face15sapinsapin /kumu-livestream-rawgated halo-livestream-raw Unsegmented source recordings behind sapinsapin/halo-livestream — the archival input to the pipeline, not a training set. 1 recording(s) · 26:56 · 9.0MB 🔒 Gated on purpose Full-length conversation between named, identifiable speakers is a very different privacy proposition from the short segments in the processed dataset, so access here is gated: request it and agree to the terms above. If what you want is segmented, quality-scored… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/kumu-livestream-raw.automatic-speech-recognitionn<1K0 likes21 downloads11d agoHugging Face16freococo /3hr_myanmar_asr_raw_audio 📚 3-Hour Burmese Speech Dataset from FOEIM Academy (ASR-ready) This is a curated ~3-hour dataset of Burmese-language audio-transcript pairs derived from the official public-service educational media of FOEIM Academy, a civic platform affiliated with FOEIM.ORG. It is structured for fine-grained automatic speech recognition (ASR) training and testing.All data is aligned from timestamped subtitle files (.srt) and segmented into high-quality .mp3 mono files with aligned transcripts. ➡️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/3hr_myanmar_asr_raw_audio.audioautomatic-speech-recognition1K<n<10K0 likes17 downloads1y agoHugging Face17Reubencf /goan-konkani-raw-audio Goan Konkani Raw Audio Private raw-ingestion index. Audio objects are preserved in their original YouTube audio container in the private Reubencf/goan-konkani-raw-audio Storage Bucket. Cleaning, repeated-Mass detection, song removal, segmentation, and transcription are derived stages; the raw source objects are not overwritten. Access and reuse remain subject to the source owners’ rights and permissions. data/pilot_metadata.jsonl records checksums, durations, source URLs, and… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/goan-konkani-raw-audio.automatic-speech-recognition0 likes11 downloads2mo agoHugging Face18freococo /voa_9000h_raw_mp3 VOA 9000h Raw MP3 Raw MP3 archive of VOA Burmese radio broadcasts (2012–2025). 8,234 full-length programs, ~4,000 hours, ~138 GB. Split across 17 tar files (voa_raw_0000.tar … voa_raw_0016.tar), 500 MP3s per tar. Source: freococo/9000hours_voa_burmese_audio — filtered to live URLs. Derived datasets: voa_myanmar_voices — 20s FLAC chunks + transcripts (498 GB) myanmar_asr — ASR model trained on this audio License Public domain (VOA staff recordings, U.S. 17 U.S.C.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_9000h_raw_mp3.audioautomatic-speech-recognition1K<n<10K0 likes10 downloads1d agoHugging Face19amn-raw /nigerian-asr-datasetgated NaijaVox ASR Dataset An ASR dataset for four Nigerian languages — Hausa, Igbo, Yoruba, and Nigerian Pidgin — sourced from Google's Waxal corpus. Built to train and evaluate automatic speech recognition models for Nigerian languages. Configs Config Language Train Validation Test Total ha Hausa 1655 296 20 1971 ig Igbo 1604 287 20 1911 yo Yoruba 2192 392 27 2611 pcm Nigerian Pidgin 1674 299 20 1993 Splits: train 84% / validation 15% / test 1%… See the full description on the dataset page: https://huggingface.co/datasets/amn-raw/nigerian-asr-dataset.audioautomatic-speech-recognition1K<n<10K1 likes9 downloads2mo agoHugging Face20hudsongouge /Podcast-Transcripts-Rawgated Podcast Transcripts Speaker-diarized transcripts from ~100k+ podcast episodes and YouTube videos (~45k hours of audio). This dataset will be gated. Only people who are part of our team may access. Splits Config Rows Description shows 124 Channels / podcast feeds (name, description, hosts, links) episodes 102,374 Episode/video metadata (title, description, guests, tags, dates) transcripts 102,374 ASR transcripts with diarized segments + YouTube… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Raw.tabularautomatic-speech-recognition100K<n<1M1 likes4 downloads3mo agoHugging Face21somu9 /raw-emoceangated raw-emocean Large-scale English speech dataset for text-to-speech (TTS) model training. Designed for autoregressive TTS architectures (TADA, CSM, VALL-E style models). Dataset Summary Metric Value Parquet shards 7 Segment duration 3–8 seconds Sample rate 24,000 Hz (mono) ASR engine NVIDIA Parakeet TDT 0.6B v3 Format Parquet with embedded audio Dataset Schema Column Type Description audio Audio Waveform array + sampling… See the full description on the dataset page: https://huggingface.co/datasets/somu9/raw-emocean.audiotext-to-speech10K<n<100K1 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.