CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01masuidrive /cv-corpus-17.0-zh-TW-client_id-grouped cv-corpus-17.0-zh-TW-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-TW-client_id-grouped.audioautomatic-speech-recognition10K<n<100K1 likes337 downloads2y agoHugging Face02masuidrive /cv-corpus-1.0-en-client_id-grouped cv-corpus-1.0-en-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 60 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-1.0-en-client_id-grouped.audioautomatic-speech-recognition100K<n<1M1 likes312 downloads2y agoHugging Face03masuidrive /cv-corpus-17.0-zh-CN-client_id-grouped cv-corpus-17.0-zh-CN-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-CN-client_id-grouped.audioautomatic-speech-recognition100K<n<1M3 likes293 downloads2y agoHugging Face04masuidrive /cv-corpus-17.0-ja-client_id-grouped cv-corpus-17.0-ja-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-ja-client_id-grouped.audioautomatic-speech-recognition10K<n<100K2 likes161 downloads2y agoHugging Face05CLiC-UB /rapnic-examplegated RAPNIC Dataset (example) Dataset Description This is an example of the full dataset, yet to be published, with 10 audio examples for 72 speakers. RAPNIC (Reconeixement Automàtic de la Parla No Intel·ligible en Català) is a Catalan speech corpus collected from individuals with speech disorders, specifically cerebral palsy and Down syndrome. This dataset was collected to develop and improve automatic speech recognition (ASR) systems that are accessible to people with speech… See the full description on the dataset page: https://huggingface.co/datasets/CLiC-UB/rapnic-example.audioautomatic-speech-recognition1K<n<10K0 likes15 downloads5mo agoHugging Face06anchpop /yap-movie-clipsgated Yap Movie Clips 164,882 short video clips of film dialogue from 366 films, one sentence per clip, in 12 languages. Every clip comes with the sentence as written in the film's official subtitles, word-level timings, the surrounding subtitle cues, an independent speech-to-text transcript of the same audio, and the phoneme sequence the sentence was expected to contain versus what a phoneme recogniser actually heard. This is the corpus behind the listening and pronunciation cards at… See the full description on the dataset page: https://huggingface.co/datasets/anchpop/yap-movie-clips.tabularautomatic-speech-recognition100K<n<1M1 likes14 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.