CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sarahwei /Taiwanese-Minnan-Example-Sentences Taiwanese Minnan Example Sentences The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems. Dataset Features Source: Ministry of Education, Taiwan (Sutian Resource Center) Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.audioautomatic-speech-recognition10K<n<100K12 likes301 downloads2y agoHugging Face02BoburAmirov /example Dataset Card for [Dataset Name] Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/example.automatic-speech-recognition0 likes107 downloads3y agoHugging Face03miracleyin /example_mmdata_mnbvc mnbvc mm dataset v2.1 MNBVC 多模态语料数据格式。原链接:https://huggingface.co/datasets/wanng/example_mmdata_mnbvc 参考实现:mm_template_mnbvc 的 mmdata_block.BLOCK_SCHEMA。schema 以那份代码为准,这个数据集是它的示例产物。 字段 字段名称 类型 字段说明 可选 实体ID string 数据的唯一标识符。用于在数据集中确定是哪一条数据。在单个数据集中确定一条数据的实体对象。 必选 md5 string 内容的 md5,用于去重与完整性校验 必选 块ID int32 一个实体对象内的标识符。用于确定一条数据内的一个部分数据。parquet 行的最小单元。 必选 块类型 string 用于保存块的类别。类别的含义为「模态」。取值见下 必选 扩展字段 string 用于保存块的元信息。为可以被成功 load 的 json 字符串。后期可继续扩展 必选… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/example_mmdata_mnbvc.audioimage-to-textn<1K2 likes57 downloads27d agoHugging Face04WhissleAI /speech-simulated-medical-examsgated Speech Simulated Medical Exams Simulated patient-physician medical exam conversations with rich speech metadata annotations. Built for training single-step ASR models that transcribe and annotate multiple concepts simultaneously, including speaker changes, emotions, intents, and roles. Dataset Details Property Value Examples 25,706 Language English Audio 16 kHz WAV Source Simulated medical interviews (respiratory focus) Features… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/speech-simulated-medical-exams.audioautomatic-speech-recognition10K<n<100K4 likes16 downloads4mo agoHugging Face05CLiC-UB /rapnic-examplegated RAPNIC Dataset (example) Dataset Description This is an example of the full dataset, yet to be published, with 10 audio examples for 72 speakers. RAPNIC (Reconeixement Automàtic de la Parla No Intel·ligible en Català) is a Catalan speech corpus collected from individuals with speech disorders, specifically cerebral palsy and Down syndrome. This dataset was collected to develop and improve automatic speech recognition (ASR) systems that are accessible to people with speech… See the full description on the dataset page: https://huggingface.co/datasets/CLiC-UB/rapnic-example.audioautomatic-speech-recognition1K<n<10K0 likes15 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.