datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Taiwanese-Minnan-Sutiau
Taiwanese-Minnan-Sutiau Dataset
The dataset consists of a curated collection of words that resemble tokens in Taiwanese Minnan (Taiwanese Hokkien), aimed at enhancing the recognition and processing of the language for various applications. Sourced from the Ministry of Education in Taiwan, this dataset serves as a valuable linguistic resource for researchers and developers engaged in language processing and recognition tasks.
Dataset Features
Source: Ministry of Education, Taiwan… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Sutiau.Taiwanese-Minnan-Example-Sentences
Taiwanese Minnan Example Sentences
The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems.
Dataset Features
Source: Ministry of Education, Taiwan (Sutian Resource Center)
Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.Taiwanese-Chinese_characters-POJ-Collectionannotations_creators:
expert-generated
language:
zh
en
language_creators:
expert-generated
license:
mit
multi-linguality:
monolingual
pretty_name: '
Taiwanese text dataset: a Chinese characters and POJ collection'
size_categories:
1M<n<10M
source_datasets:
original
tags: []
task_categories:
text-classification
feature-extraction
task_ids:
multi-label-classification
multi-class-classification
taiwanese_english_translationThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.Taiwanese_ASRcv-taiwaneseErrorSpanAnnotation-for-Taiwanese-Hokkien
Error Span Annotation for Taiwanese Hokkien
The Taiwanese Hokkien subset of the SiniticMTError benchmark (Liu et al., 2026).
Human-annotated machine-translation error-span evaluation data for the
Mandarin → Taiwanese Hokkien (Tâi-gí) direction. Each instance contains a
Mandarin source sentence, a Taiwanese Hokkien machine translation, a reference
translation, and expert error-span annotations with severity labels and a
segment-level quality score.
Language pair: Mandarin (zh) →… See the full description on the dataset page: https://huggingface.co/datasets/350016z/ErrorSpanAnnotation-for-Taiwanese-Hokkien.taiwanese-youtube-ocr-robusttaiwanese-hokkien-ocr-testtaiwanese-youtube-ocr-first-5ima-taiwanese-corpus-merged
IMA Taiwanese Corpus(台語語料總集)
本資料集為台語語料總集,目的在於將原先分散於多位作者 dataset repo 的台語語料統一整併,
提供「一次申請、持續更新」的存取方式。
目錄結構
所有來源資料皆保留於 data/ 之下,每個子資料夾對應一個原始來源 repo,例如:
data/taigi-literature-ots
data/taigi-literature-tks
data/taigi-literature-abt
(其餘同理)
每個子資料夾內保留原始 README 與語料檔案,以利來源追溯與審核。
使用方式
申請者只需申請本 dataset(本 repo)一次,即可取得所有台語語料。
未來新增作者或新增語料將直接更新於本 repo,不需重複申請。
Taiwanese-Chinese_characters-POJ-Collectiontaiwanese-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Taiwan
The Synthetic Taiwan Passports Dataset compiles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Each record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/taiwanese-passports.taiwanese-college-studentsAnonymous survey responses by Taiwanese college students.
Dolphin_Model_Chinese-Taiwanese-Speech-Recognition-Corpus
ID
King-ASR-044
Duration
2181.7 hours
Description
This dataset was recorded in both quiet and noisy environments, with 5226 speakers participating, including 2352 males and 2874 females. All speakers involved in the recording were professionally selected to ensure standard pronunciation and clear enunciation. The recorded text covers information such as news, daily conversations, and SMS messages.
URL… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Dolphin_Model_Chinese-Taiwanese-Speech-Recognition-Corpus.taiwanese-youtube-ocr-testTaiwanese_Hokkien_Tone_RecognitionTaiwanese-dataset-v0
