CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01michaelcacioli /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes308 downloads3mo agoHugging Face02Jaspernl /The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden" Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io Dataset Summary The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.audioautomatic-speech-recognition10K<n<100K1 likes156 downloads2y agoHugging Face03flagship-ai /ghomala-spoken-bible Ghomálá' Spoken New Testament — aligned audio + trilingual text Part of the Lingo / NativeAI language-preservation project. This is ~20 hours of spoken Ghomálá' (Ghomala, ISO bbj; a Grassfields Bantu language of West Cameroon) — recorded readings of the New Testament — aligned chapter-by-chapter with parallel text in Ghomálá', French, and English. Spoken-language data is exactly what oral-first Cameroonian languages lack, which makes this a rare resource for building ASR, TTS… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/ghomala-spoken-bible.audioautomatic-speech-recognition10K<n<100K0 likes123 downloads4mo agoHugging Face04vnahata /SpokenWikipedia-retrieval Spoken Wikipedia speech-text retrieval (MTEB) Volunteer readings of Wikipedia articles paired with the article lead, in Dutch, English, German, Spanish and French. Recordings come from Wikimedia Commons, which is free by site policy, and the lead text from each Wikipedia, which is CC-BY-SA. The set is published as cc-by-sa-4.0. Only the first 60 seconds of each reading is kept, since readers start at the lead. One recording per article. Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/SpokenWikipedia-retrieval.audioautomatic-speech-recognitionn<1K0 likes113 downloads23d agoHugging Face05ymoslem /SpokenWords-GA-EN-MTed Dataset Card for Dataset Name This is the Irish portion of the Spoken Words dataset (available at MLCommons/ml_spoken_words), with merged splits “train”, “validation”, and “test”, augmented with machine translation. The Irish sentences are automatically translated into English using Google Translation API. The dataset includes approximately 3 hours and 2 minutes of audio (03:02:02), spoken by multiple narrators. Dataset Structure Dataset({ features: ['keyword'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/SpokenWords-GA-EN-MTed.audioautomatic-speech-recognition10K<n<100K1 likes64 downloads2y agoHugging Face06TigreGotico /SpokenPortugueseGeographicalSocialVarieties Spoken Portuguese - Geographical and Social Varieties dataset source: https://www.clul.ulisboa.pt (1995-1997 - European Commission DGXXII, Programme LINGUA/SOCRATES) The project is concluded and the materials are published in CD-ROM, with the exclusive publishing support of Instituto Camões, under the title Português Falado - Documentos Autênticos: Gravações áudio com transcrição alinhada. Its distribution outside of Portugal is ensured by Instituto Camões and in Portugal by CLUL.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/SpokenPortugueseGeographicalSocialVarieties.audioautomatic-speech-recognitionn<1K0 likes57 downloads1y agoHugging Face07Jarbas /SpokenPortugueseGeographicalSocialVarieties_splitssentence splits from SpokenPortugueseGeographicalSocialVarieties generated via forced alignment audioautomatic-speech-recognition1K<n<10K0 likes53 downloads1y agoHugging Face08Tohirju /tajik-spoken-instructionsgated Tajik Spoken Instructions (Q&A) Synthetic Tajik speech of dictionary and language-exercise questions, each paired with its written answer — spoken instruction in, text answer out. 648,842 clips · ~540 hours · 2 voices (male + female) Questions cover word meanings, antonyms, etymology, usage and grammar Numbers expanded to spoken Tajik Deduplicated; Latin-script and mis-encoded rows removed Contents file what audio_*.tar the wav files manifest.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/tajik-spoken-instructions.text-to-speech100K<n<1M0 likes40 downloads17d agoHugging Face09Tohirju /tajik-spoken-questions-omnigated Tajik Spoken Questions — Omni-a3 35,068 clips · 32.62 hours · Tajik (tg) · 24 kHz mono ⚠️ This audio is SYNTHETIC (TTS-generated), not real speech Every question clip here was generated by a text-to-speech model, not recorded from a human. It is deliberately published as a separate repo from Tohirju/tajik-audio and Tohirju/tajik-asr-full-data, which contain real recorded speech — do not mix them without knowing which is which. Do not use this set to train or… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/tajik-spoken-questions-omni.automatic-speech-recognition0 likes28 downloads21d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.