CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lingamvamshikrishnareddy /ramanv-tts-all-rawgated ramanv-tts-all-raw Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages. textautomatic-speech-recognition1M<n<10M0 likes4.4k downloads12d agoHugging Face02Whispering-GPT /linustechtips-transcript-audio Dataset Card for "linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips. Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.audioautomatic-speech-recognitionn<1K4 likes2.3k downloads4y agoHugging Face03linagora /SUMM-RENote: if the data viewer is not working, use the "example" subset. SUMM-RE The SUMM-RE dataset is a collection of transcripts of French conversations, aligned with the audio signal. It is a corpus of meeting-style conversations in French created for the purpose of the SUMM-RE project (ANR-20-CE23-0017). The full dataset is described in Hunter et al. (2024): "SUMM-RE: A corpus of French meeting-style conversations". Created by: Recording and manual correction of the corpus was… See the full description on the dataset page: https://huggingface.co/datasets/linagora/SUMM-RE.audioautomatic-speech-recognitionn<1K5 likes2.2k downloads2y agoHugging Face04linagora /linto-dataset-audio-ar-tn LinTO DataSet Audio for Arabic Tunisian A collection of Tunisian dialect audio and its annotations for STT task This is the first packaged version of the datasets used to train the Linto Tunisian dialect with code-switching STT (linagora/linto-asr-ar-tn). Dataset Summary Dataset composition Sources Data Table Data sources Content Types Languages and Dialects Example use (python) License Citations Dataset Summary The LinTO DataSet Audio for Arabic Tunisian is a diverse… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn.audioautomatic-speech-recognition10K<n<100K22 likes1.5k downloads1y agoHugging Face05linhtran92 /viet_bud500gated Bud500: A Comprehensive Vietnamese ASR Dataset Introducing Bud500, a diverse Vietnamese speech corpus designed to support ASR research community. With aprroximately 500 hours of audio, it covers a broad spectrum of topics including podcast, travel, book, food, and so on, while spanning accents from Vietnam's North, South, and Central regions. Derived from free public audio resources, this publicly accessible dataset is designed to significantly enhance the work of developers and… See the full description on the dataset page: https://huggingface.co/datasets/linhtran92/viet_bud500.audioautomatic-speech-recognition100K<n<1M73 likes703 downloads3y agoHugging Face06linagora /linto-dataset-audio-ar-tn-augmented LinTO DataSet Audio for Arabic Tunisian Augmented A collection of Tunisian dialect audio and its annotations for STT task This is the augmented datasets used to train the Linto Tunisian dialect with code-switching STT linagora/linto-asr-ar-tn. Dataset Summary Dataset composition Sources Content Types Languages and Dialects Example use (python) License Citations Dataset Summary The LinTO DataSet Audio for Arabic Tunisian Augmented is a dataset that builds on LinTO… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn-augmented.audioautomatic-speech-recognition100K<n<1M7 likes675 downloads1y agoHugging Face07vnahata /LinguaLibre-word-retrieval Lingua Libre spoken-word retrieval (MTEB) Single words read aloud by volunteers, paired with the written word, across many languages. Recordings come from Lingua Libre, a Wikimedia project, hosted on Wikimedia Commons, which is free by site policy. Published as cc-by-sa-4.0. Audio is 16 kHz Opus. Bare punctuation and read sentences are excluded, and each word is kept once. Built by scripts/data/lingua_libre/create_data.py in the MTEB repo. audioautomatic-speech-recognition1K<n<10K1 likes488 downloads23d agoHugging Face08Congo-digital-service /audios-lingala-annotatees Annotated Lingala Dataset – Full Version Description This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models. It includes: the original audio files (viewable directly in the Hugging Face viewer) text transcriptions Mel spectrograms tokenized labels Overall statistics Metric Value Total volume 5 h 0 min 18 s Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.audioautomatic-speech-recognition10K<n<100K0 likes322 downloads18d agoHugging Face09KasuleTrevor /Lingala_100hrs Lingala 100hrs 110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated from three publicly available CC-BY-4.0 corpora for ASR research. Composition Counts from a full-pass audit on 2026-07-09: Source Upstream location Rows Splits AfriVoice (Lingala) https://huggingface.co/datasets/DigitalUmuganda/AfriVoice 17,544 train (16,144), validation (915), test (485) LRSC (Lingala Read Speech Corpus)… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs.audioautomatic-speech-recognition10K<n<100K0 likes310 downloads3mo agoHugging Face10Congo-digital-service /audios-lingala-annotatees-v2 Annotated Lingala Audio — canonical corpus Annotated Lingala speech for open automatic speech recognition research and for fine-tuning speech models. This release is a full reconstruction of the corpus from its source recordings and annotations. It supersedes Congo-digital-service/audios-lingala-annotatees, which is deprecated — see Relationship to the previous release below. What this dataset contains Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.audioautomatic-speech-recognition10K<n<100K0 likes169 downloads14d agoHugging Face11cublya /jam-alt-lines Jam-ALT Lines Jam-ALT Lines is a line-level version of the Jam-ALT lyrics transcription dataset. Unlike Jam-ALT, this dataset contains one audio segment for each lyrics line, facilitating research that considers each line as a separate unit. [!tip] See the Jam-ALT project website for details and the JamendoLyrics community for related datasets. Dataset flavors Lyrics lines may overlap in time, which makes it impossible to have a one-to-one correspondence between… See the full description on the dataset page: https://huggingface.co/datasets/cublya/jam-alt-lines.audioautomatic-speech-recognition10K<n<100K0 likes150 downloads9mo agoHugging Face12linq1005 /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/linq1005/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes101 downloads2mo agoHugging Face13lindonghello /omnilingual-asr-corpus Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn", `glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/lindonghello/omnilingual-asr-corpus.audioautomatic-speech-recognition100K<n<1M0 likes96 downloads5mo agoHugging Face14jamendolyrics /jam-alt-lines Jam-ALT Lines Jam-ALT Lines is a line-level version of the Jam-ALT lyrics transcription dataset. Unlike Jam-ALT, this dataset contains one audio segment for each lyrics line, facilitating research that considers each line as a separate unit. [!tip] See the Jam-ALT project website for details and the JamendoLyrics community for related datasets. Dataset flavors Lyrics lines may overlap in time, which makes it impossible to have a one-to-one correspondence between… See the full description on the dataset page: https://huggingface.co/datasets/jamendolyrics/jam-alt-lines.audioautomatic-speech-recognition10K<n<100K2 likes94 downloads1y agoHugging Face15Bretagne /Lingua_Libre_br [!NOTE] Dataset origin: https://lingualibre.org/LanguagesGallery/ Description 1h 0min 2s d'audio.Les fichiers .ogg ont été convertis en .mp3.L'extraction date de janvier 2026. audioautomatic-speech-recognition1K<n<10K0 likes49 downloads2mo agoHugging Face16fosters /astryd_lindgren_braty_lvinae_sertsa_all Браты Ільвінае сэрца Аўтар / Author: Астрыд ЛіндгрэнМова / Language: Беларуская (Belarusian) Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд. Частка калекцыі Belarusian Audiobooks (native). Радкоў у датасеце 2,116 Працягласць 6 гадз 31 хв Частата дыскрэтызацыі 44100 Hz Каналы мона Даўжыня фрагмента да 30 с Структура Кожны радок змяшчае: audio — аўдыёфрагмент (native SR, мона… See the full description on the dataset page: https://huggingface.co/datasets/fosters/astryd_lindgren_braty_lvinae_sertsa_all.audioautomatic-speech-recognition1K<n<10K0 likes40 downloads3mo agoHugging Face17Whispering-GPT /whisper-transcripts-linustechtips Dataset Card for "whisper-transcripts-linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the channel. channel_id: Id of the youtube channel. title: Title given to… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/whisper-transcripts-linustechtips.textautomatic-speech-recognition1K<n<10K2 likes35 downloads4y agoHugging Face18BantuLanguagesInitiative /lingala_real_eval_benchmark_croped Lingala Real Eval Benchmark Cropped Small cropped real-world Lingala audio benchmark for testing BLI ASR 0. The dataset contains short audio clips cropped from longer real-world files, covering different domains such as news, catechesis, comedy, cartoon and interview speech. This dataset is intended for quick qualitative ASR testing and human review. It is not a training dataset. audioautomatic-speech-recognitionn<1K1 likes26 downloads4mo agoHugging Face19Speech-data /Lingala-Speech-Dataset Lingala Dataset Metadata Field Value 📜 License CC BY-NC-ND 4.0 🎯 Task Categories Automatic Speech Recognition 🌍 Language Lingala (ln) 🏷️ Tags Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Lingala 📦 Size Category n < 1K audioautomatic-speech-recognitionn<1K0 likes14 downloads6mo agoHugging Face20fosters /astryd_lindgren_braty_lvinae_sertsa_output_original Браты Ільвінае сэрца — арыгінальнае аўдыё Аўтар / Author: Астрыд ЛіндгрэнМова / Language: Беларуская (Belarusian) Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці. Частка калекцыі Ministerskija — корпус беларускіх аўдыёкніг. Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя): astryd_lindgren_braty_lvinae_sertsa_output Доўгасць аўдыё 7h06m Радкоў у датасеце 1,835 Структура Кожны радок змяшчае: audio — арыгінальны… See the full description on the dataset page: https://huggingface.co/datasets/fosters/astryd_lindgren_braty_lvinae_sertsa_output_original.audioautomatic-speech-recognition1K<n<10K0 likes14 downloads4mo agoHugging Face21slaycreep /jam-alt-lines Jam-ALT Lines Jam-ALT Lines is a line-level version of the Jam-ALT lyrics transcription dataset. Unlike Jam-ALT, this dataset contains one audio segment for each lyrics line, facilitating research that considers each line as a separate unit. [!tip] See the Jam-ALT project website for details and the JamendoLyrics community for related datasets. Dataset flavors Lyrics lines may overlap in time, which makes it impossible to have a one-to-one correspondence between… See the full description on the dataset page: https://huggingface.co/datasets/slaycreep/jam-alt-lines.audioautomatic-speech-recognition10K<n<100K0 likes9 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.