CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01michaelcacioli /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes308 downloads3mo agoHugging Face02Jaspernl /The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden" Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io Dataset Summary The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.audioautomatic-speech-recognition10K<n<100K1 likes156 downloads2y agoHugging Face03flagship-ai /ghomala-spoken-bible Ghomálá' Spoken New Testament — aligned audio + trilingual text Part of the Lingo / NativeAI language-preservation project. This is ~20 hours of spoken Ghomálá' (Ghomala, ISO bbj; a Grassfields Bantu language of West Cameroon) — recorded readings of the New Testament — aligned chapter-by-chapter with parallel text in Ghomálá', French, and English. Spoken-language data is exactly what oral-first Cameroonian languages lack, which makes this a rare resource for building ASR, TTS… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/ghomala-spoken-bible.audioautomatic-speech-recognition10K<n<100K0 likes123 downloads4mo agoHugging Face04vnahata /SpokenWikipedia-retrieval Spoken Wikipedia speech-text retrieval (MTEB) Volunteer readings of Wikipedia articles paired with the article lead, in Dutch, English, German, Spanish and French. Recordings come from Wikimedia Commons, which is free by site policy, and the lead text from each Wikipedia, which is CC-BY-SA. The set is published as cc-by-sa-4.0. Only the first 60 seconds of each reading is kept, since readers start at the lead. One recording per article. Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/SpokenWikipedia-retrieval.audioautomatic-speech-recognitionn<1K0 likes113 downloads24d agoHugging Face05ymoslem /SpokenWords-GA-EN-MTed Dataset Card for Dataset Name This is the Irish portion of the Spoken Words dataset (available at MLCommons/ml_spoken_words), with merged splits “train”, “validation”, and “test”, augmented with machine translation. The Irish sentences are automatically translated into English using Google Translation API. The dataset includes approximately 3 hours and 2 minutes of audio (03:02:02), spoken by multiple narrators. Dataset Structure Dataset({ features: ['keyword'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/SpokenWords-GA-EN-MTed.audioautomatic-speech-recognition10K<n<100K1 likes64 downloads2y agoHugging Face06TigreGotico /SpokenPortugueseGeographicalSocialVarieties Spoken Portuguese - Geographical and Social Varieties dataset source: https://www.clul.ulisboa.pt (1995-1997 - European Commission DGXXII, Programme LINGUA/SOCRATES) The project is concluded and the materials are published in CD-ROM, with the exclusive publishing support of Instituto Camões, under the title Português Falado - Documentos Autênticos: Gravações áudio com transcrição alinhada. Its distribution outside of Portugal is ensured by Instituto Camões and in Portugal by CLUL.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/SpokenPortugueseGeographicalSocialVarieties.audioautomatic-speech-recognitionn<1K0 likes57 downloads1y agoHugging Face07Jarbas /SpokenPortugueseGeographicalSocialVarieties_splitssentence splits from SpokenPortugueseGeographicalSocialVarieties generated via forced alignment audioautomatic-speech-recognition1K<n<10K0 likes53 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.