CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sapinsapin /filipinospeechcorpus Filipino Speech Corpus (FSC) Studio-recorded Filipino read, spontaneous, and word-level speech — 125 speakers, packaged as ready-to-stream Parquet. 313,322 transcribed segments · 65.1 hours · 125 speakers · 16kHz mono This is the Filipino Speech Corpus (Sagum), recorded in a controlled setting and hand/machine transcribed with Transcriber. This repo repackages the original .wav + .trs volumes as segment-level Parquet with inline audio, so you can stream it without… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/filipinospeechcorpus.audioautomatic-speech-recognition100K<n<1M3 likes609 downloads1mo agoHugging Face02sapinsapin /pld Philippine Language Dataset (PLD) Ten Philippine languages, 980 speakers, 448 hours of prompted speech — one of the largest multilingual Philippine speech collections available as Parquet. 334,268 utterances · 448.2 hours · 980 speakers · 10 languages · 16kHz mono ▶ Try the models in your browser — transcribe, synthesize, or convert a voice in any of the ten languages, from your microphone or the preloaded clips. Collected by the University of the Philippines Diliman… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/pld.audioautomatic-speech-recognition100K<n<1M0 likes475 downloads16h agoHugging Face03sapinsapin /kumu-livestream-segmentedgated halo-livestream Real Taglish code-switching from livestreams — every segment carries forced-alignment confidence, ASR round-trip CER, SNR, loudness and overlap flags. 62 segments · 3 speakers · seed release 🌱 This is a seed release — 62 segments, about 7 minutes It exists to publish the pipeline and the schema, not to be a training corpus. Nothing here is big enough to train on. What is worth your time is the per-segment quality metadata below — and the… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/kumu-livestream-segmented.audioautomatic-speech-recognitionn<1K0 likes15 downloads1mo agoHugging Face04breno30 /Sapinhaaudion<1K0 likes6 downloads3y agoHugging Face05NbAiLab /nrk-sapmi-podcastsaudion<1K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.