CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01carlosdanielhernandezmena /ravnursson_asr Dataset Card for ravnursson_asr Dataset Summary The corpus "RAVNURSSON FAROESE SPEECH AND TRANSCRIPTS" (or RAVNURSSON Corpus for short) is a collection of speech recordings with transcriptions intended for Automatic Speech Recognition (ASR) applications in the language that is spoken at the Faroe Islands (Faroese). It was curated at the Reykjavík University (RU) in 2022. The RAVNURSSON Corpus is an extract of the "Basic Language Resource Kit 1.0" (BLARK 1.0) [1] developed… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/ravnursson_asr.audioautomatic-speech-recognition10K<n<100K3 likes955 downloads1y agoHugging Face02Data-Science-Nigeria /voice-of-care-health-dataset Voice of Care AI for Global Health Benchmark Dataset Overview This dataset contains spoken Hausa Health datasets with rich annotations covering emotion, intent, speaker demographics, and dialect variation, intended for speech and NLP research. Dataset Summary Property Details Language Hausa Modality Audio + Text Task(s) e.g. Speech Recognition, Emotion Detection, Dialect Identification Version 1.0.0 🛠️ Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Data-Science-Nigeria/voice-of-care-health-dataset.audioautomatic-speech-recognition10K<n<100K1 likes84 downloads2mo agoHugging Face03kawshikbuet17 /bengali-telecom-customer-care-speech-v2 Bengali Telecom Customer Care Synthetic Speech Dataset v2 Dataset Description This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts. The dataset is intended for experiments with: Bengali ASR/STT Bengali TTS Speech-to-text preprocessing Telecom/customer-care domain adaptation Synthetic speech research This is a second version of the Bengali Telecom Customer Care Synthetic Speech Dataset. It follows the same… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech-v2.audiotext-to-speech1K<n<10K0 likes72 downloads3mo agoHugging Face04carlosdanielhernandezmena /toy_corpus_asr_caThis is an example of a repository with a standard data loader. The audio files are compressed in tar format. Since this repository contains very few audio files, it can be used to test certain scripts in local machines. audioautomatic-speech-recognitionn<1K0 likes24 downloads2y agoHugging Face05kawshikbuet17 /bengali-telecom-customer-care-speech Bengali Telecom Customer Care Synthetic Speech Dataset Dataset Description This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts. The dataset is intended for experiments with: Bengali ASR/STT Bengali TTS Speech-to-text preprocessing Telecom/customer-care domain adaptation Synthetic speech research Important Disclosure This is a synthetic speech dataset generated using the OmniVoice TTS system in… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech.audiotext-to-speech10K<n<100K0 likes23 downloads4mo agoHugging Face06carlosdanielhernandezmena /toy_corpus_asr_esThis is an example of a repository with a standard data loader. The audio files are compressed in tar format. Since this repository contains very few audio files, it can be used to test certain scripts in local machines. audioautomatic-speech-recognitionn<1K0 likes20 downloads2y agoHugging Face07carlosdanielhernandezmena /prueba_parquetThis is an example of a repository with parquet files only. audioautomatic-speech-recognitionn<1K0 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.