CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sarahwei /Taiwanese-Minnan-Sutiau Taiwanese-Minnan-Sutiau Dataset The dataset consists of a curated collection of words that resemble tokens in Taiwanese Minnan (Taiwanese Hokkien), aimed at enhancing the recognition and processing of the language for various applications. Sourced from the Ministry of Education in Taiwan, this dataset serves as a valuable linguistic resource for researchers and developers engaged in language processing and recognition tasks. Dataset Features Source: Ministry of Education, Taiwan… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Sutiau.audioautomatic-speech-recognition10K<n<100K19 likes420 downloads2y agoHugging Face02sarahwei /Taiwanese-Minnan-Example-Sentences Taiwanese Minnan Example Sentences The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems. Dataset Features Source: Ministry of Education, Taiwan (Sutian Resource Center) Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.audioautomatic-speech-recognition10K<n<100K12 likes309 downloads2y agoHugging Face03gacky1601 /Taiwanese_ASRaudio1K<n<10K2 likes52 downloads2y agoHugging Face04quyanh /cv-taiwaneseaudio1K<n<10K0 likes23 downloads2mo agoHugging Face05350016z /ErrorSpanAnnotation-for-Taiwanese-Hokkien Error Span Annotation for Taiwanese Hokkien The Taiwanese Hokkien subset of the SiniticMTError benchmark (Liu et al., 2026). Human-annotated machine-translation error-span evaluation data for the Mandarin → Taiwanese Hokkien (Tâi-gí) direction. Each instance contains a Mandarin source sentence, a Taiwanese Hokkien machine translation, a reference translation, and expert error-span annotations with severity labels and a segment-level quality score. Language pair: Mandarin (zh) →… See the full description on the dataset page: https://huggingface.co/datasets/350016z/ErrorSpanAnnotation-for-Taiwanese-Hokkien.tabulartranslationn<1K0 likes17 downloads2mo agoHugging Face06alexachang /taiwanese-youtube-ocr-robustaudion<1K0 likes15 downloads1y agoHugging Face07alexachang /taiwanese-hokkien-ocr-testaudion<1K0 likes13 downloads1y agoHugging Face08alexachang /taiwanese-youtube-ocr-first-5audion<1K0 likes13 downloads1y agoHugging Face09ud-synthetic /taiwanese-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - Taiwan The Synthetic Taiwan Passports Dataset compiles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Each record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/taiwanese-passports.textimage-to-textn<1K1 likes8 downloads2mo agoHugging Face10alexachang /taiwanese-youtube-ocr-testaudion<1K0 likes5 downloads1y agoHugging Face11DynamicSuperb /Taiwanese_Hokkien_Tone_Recognitiongatedaudion<1K1 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.