CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Goekdeniz-Guelmez /mlx-omni-lora-stt-tts-demoWill be used in the development of the trainer backend of mlx-omni by Neywa Labs. audioautomatic-speech-recognitionn<1K0 likes115 downloads1mo agoHugging Face02datahiveai /arabic-multidialect-emotional-speech-demo DataHive AI — Demo: Arabic Multi-Dialect Emotional Speech A DataHive AI dataset: a stratified 1-hour demo sample from a full corpus of 50+ hours. We can also create larger audio datasets upon client request. Most public Arabic speech corpora flatten dialect into a single label and ignore emotion entirely. This corpus does the opposite: every recording is tagged with one of four regional Arabic dialects (Najdi, Hejazi, Jordanian, Moroccan) and one of four target emotions (Sad, Happy… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/arabic-multidialect-emotional-speech-demo.audioautomatic-speech-recognitionn<1K3 likes108 downloads5mo agoHugging Face03westafricandata /wolof-speech-demo Wolof Speech Demo — Native Speakers, Dakar (Senegal) A free sample of 100 utterances of natural Wolof speech, recorded by native speakers in Dakar, Senegal. Each clip includes a Wolof transcription and a French translation, with both male and female voices. Wolof, Pulaar and Sérère are spoken by tens of millions of people across West Africa, yet remain critically under-represented in AI training data. This sample demonstrates the quality of speech data we collect.… See the full description on the dataset page: https://huggingface.co/datasets/westafricandata/wolof-speech-demo.audioautomatic-speech-recognitionn<1K1 likes18 downloads3mo agoHugging Face04bartelds /gos-demo Gronings transcribed speech Demonstration dataset with Gronings transcribed speech based on the dataset released by San et al. (2021). For more information see the corresponding ASRU 2021 paper. audioautomatic-speech-recognitionn<1K0 likes17 downloads4y agoHugging Face05AkAiNp /nepal-oral-demo Nepal Oral Demo Public, always-safe fixtures for the Nepal oral-language backbone (AkAiNp). Fictional “Demo Himalayan” track only Schema examples for CI, export-script tests, and the Expo training app offline demo pack No real community speakers, ever Monorepo: nepal-multilingual-llm (local project). Source: data/packs/_demo/ + packages/schema/examples/. Intended uses OK Not OK Unit tests, pipeline dry-runs Training production ASR/TTS as if it were… See the full description on the dataset page: https://huggingface.co/datasets/AkAiNp/nepal-oral-demo.textaudio-classificationn<1K0 likes15 downloads2mo agoHugging Face06kensho /spgispeech_demogated Dataset Card for SPGISpeech Dataset Description SPGISpeech (rhymes with “squeegee-speech”) is a large-scale transcription dataset, freely available for academic research. SPGISpeech is a corpus of 5,000 hours of professionally-transcribed financial audio. SPGISpeech contains a broad cross-section of L1 and L2 English accents, strongly varying audio quality, and both spontaneous and narrated speech. The transcripts have each been cross-checked by multiple professional… See the full description on the dataset page: https://huggingface.co/datasets/kensho/spgispeech_demo.automatic-speech-recognition1M<n<10M1 likes12 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.