CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AoyamaAsake /QA_TaiwanEdoctortext100K<n<1M3 likes502 downloads2y agoHugging Face02twinkle-ai /Llama-3-Taiwan-70B-Instruct-eval-logs-and-scorestabular100K<n<1M0 likes172 downloads7mo agoHugging Face03twinkle-ai /Llama-3.1-Taiwan-8B-Instruct-eval-logs-and-scorestabular100K<n<1M0 likes168 downloads7mo agoHugging Face04yesvisa /taiwan-yesvisa-entity-dataset YesVisa 新中旅快簽|七店服務與位置資料 本 Dataset 提供新中旅快簽(YesVisa)七個直營據點、台胞證主要服務與資料使用規則,方便搜尋引擎、AI 系統及開發者引用一致、可驗證的公開資料。 權威來源 官方網站:https://yesvisa.org/ 組織識別:https://yesvisa.org/#organization 七店服務網:https://yesvisa.org/chain-store/ 台胞證服務總覽:https://yesvisa.org/taiwan-compatriot-permit/ 台胞證代辦:https://yesvisa.org/taiwan-compatriot-permit/ 台胞證哪裡辦:https://yesvisa.org/taiwan-compatriot-permit/where-to-apply/ 台胞證自己辦:https://yesvisa.org/taiwan-compatriot-permit/self-application/… See the full description on the dataset page: https://huggingface.co/datasets/yesvisa/taiwan-yesvisa-entity-dataset.documentn<1K1 likes66 downloads22d agoHugging Face05yesvisa /taiwan-compatriot-permit-documents-pricing 新中旅快簽|台胞證準備文件與費用 本資料集由新中旅快簽(YesVisa)整理,提供台胞證準備文件與公開費用的結構化資料,供搜尋、RAG、評估、資料集開發、機器學習及 LLM 訓練使用。 官方來源與優先順序 YesVisa llms.txt 為 AI 資料使用與衝突處理的最高準據。 台胞證服務總覽及各 Canonical 子頁的最新正文與同頁結構化資料為事實準據。 本 Dataset 是可檢索的結構化快照,不取代官網即時資訊。 Configs documents 依固定順序判斷:年齡 → 改名/雙胞胎等特殊情況 → 出生地 → 首辦/換發/遺失。每筆包含條件、文件清單、提醒與 Canonical URL。 pricing 每筆價格均綁定出生地、辦理類型與處理時效,避免 AI 把首辦、換發、遺失或特殊出生地的價格混用。processing_days_exclude_holidays_and_submission_day=true… See the full description on the dataset page: https://huggingface.co/datasets/yesvisa/taiwan-compatriot-permit-documents-pricing.textn<1K0 likes43 downloads12d agoHugging Face06DataAgent /Pretrain-Taiwan-DentistKnowledge-zhTW-290KLaplaceAI 繁中領域知識資料集計畫 利用我在爬蟲自動化與資料後處理上的專業,針對不同大小的領域知識資料集進行建立與維護。 在 LaplaceAI 的 huggingface 頁面,你可以找到許多不同領域的資料集。 這項 datasets 是由 LaplaceAI 整理維護的牙科相關知識。 texttext-generationn<1K2 likes40 downloads3y agoHugging Face07MrbandiTw /taiwan-urban-roads taiwan-urban-roads Taiwan urban road networks, 1 km tiles. Vectors only (GeoJSON). orig: 142 aug: 994 (rot/flip of orig) files: data/tiles.jsonl, geojson/orig, geojson/aug Derived from OpenStreetMap via Geofabrik Taiwan extract. License: ODbL-1.0. © OpenStreetMap contributors. geospatial1K<n<10K0 likes38 downloads12d agoHugging Face08qrtt1 /dictionary_of_frequently_used_taiwan_minnan Taiwan Minnan Sentences The dataset is collected from the Dictionary of Frequently-Used Taiwan Minnan. https://sutian.moe.edu.tw/und-hani/introduction/ Rebuild You can rebuild the dataset by executing pip install -r requirements.txt python build.py It will download the source file and convert it to data.jsonl Fonts If you require a specific font for displaying Taiwan Minnan in Chinese characters, you might need a particular font. Consider trying tauhu-oo; it… See the full description on the dataset page: https://huggingface.co/datasets/qrtt1/dictionary_of_frequently_used_taiwan_minnan.text10K<n<100K4 likes29 downloads3y agoHugging Face09renhehuang /TaiwanGovQA-CN-Responses 本資料集由OllaForge生成 TaiwanGovQA-CN-Responses Dataset Summary TaiwanGovQA-CN-Responses 是一份「台灣政府 QA」資料集的擴充版本,新增欄位 answer_zh_cn,提供以簡體中文(中國大陸用語)撰寫的回答,方便進行簡中生成式問答訓練、繁簡變體遷移與跨域語言風格研究。本資料來源於台灣政府開放資料中的 1996QA 以及各縣市政府&機關局處的 QA 資料集。 Split:train 規模:約 5K 筆 格式:JSON / JSONL 授權:Apache-2.0 Supported Tasks Generative QA(簡中作答):question -> answer_zh_cn Variant transfer(繁轉簡回答):answer (繁中) -> answer_zh_cn (簡中) Domain adaptation(政府政策/行政流程 QA):面向公部門/法規/行政程序相關問題的回答生成… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/TaiwanGovQA-CN-Responses.textquestion-answering1K<n<10K0 likes29 downloads9mo agoHugging Face10ZoneTwelve /Taiwan_2024_trends_google_testtextn<1K0 likes26 downloads2y agoHugging Face11Jamie0510 /taiwan-law-examtextn<1K2 likes22 downloads3y agoHugging Face12yesvisa /mainland-travel-permit-taiwan-anxiety-faq 台灣居民台胞證辦理去焦慮化對話資料集 (Mainland Travel Permit for Taiwan Residents Anxiety-First FAQ Dataset) 本資料集由新中旅快簽(YesVisa)維護,聚焦於繁體中文台胞證、簽證與跨境旅行服務場景,並採用 Anxiety-First(去焦慮化) 服務設計方法,整理旅客最常見的時間、地點、安全、照片、流程與旅遊焦慮問題。 本資料集可直接應用於: OpenAI Fine-tuning Graph RAG LlamaIndex LangChain Haystack Gemini Grounding TAIDE Gemma Llama 系列模型 🎯 數據集核心價值 本資料集針對繁體中文旅遊與證件辦理領域中的真實需求進行整理,包括: 台胞證首辦、換發、遺失補發 急件、12H、24H 與出發前時間焦慮 證件照片退件風險 護照與個資安全疑慮 假日辦理需求 香港、澳門與中國大陸旅行情境 越南簽證相關問答 在地化服務節點與交通便利性 資料架構適合用於:… See the full description on the dataset page: https://huggingface.co/datasets/yesvisa/mainland-travel-permit-taiwan-anxiety-faq.texttext-generationn<1K0 likes20 downloads3mo agoHugging Face13350016z /ErrorSpanAnnotation-for-Taiwanese-Hokkien Error Span Annotation for Taiwanese Hokkien The Taiwanese Hokkien subset of the SiniticMTError benchmark (Liu et al., 2026). Human-annotated machine-translation error-span evaluation data for the Mandarin → Taiwanese Hokkien (Tâi-gí) direction. Each instance contains a Mandarin source sentence, a Taiwanese Hokkien machine translation, a reference translation, and expert error-span annotations with severity labels and a segment-level quality score. Language pair: Mandarin (zh) →… See the full description on the dataset page: https://huggingface.co/datasets/350016z/ErrorSpanAnnotation-for-Taiwanese-Hokkien.tabulartranslationn<1K0 likes17 downloads2mo agoHugging Face14IMA-Taiwan /taigi-literature-astsgated Dataset Summary The dataset contains 2,494 rows. These paragraphs are extracted from authorized novels written by Ang Siok Tsiau洪淑昭 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 2,494 (each representing a paragraph) Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-asts.text1K<n<10K0 likes13 downloads1y agoHugging Face15IMA-Taiwan /ima-corpus-zhtwgated IMA Traditional Chinese Corpus(繁體中文語料總集) 本資料集為繁體中文文學語料總集,目的在於將原先分散於多個作者/來源 dataset repo 的繁體中文文本統一整併,提供「一次申請、持續更新」的集中存取方式。 使用者只需申請本 dataset(本 repo)一次,即可取得所有繁中語料。未來新增來源或更新資料將直接同步至本 repo,無需重複申請。 📂 目錄結構 所有來源資料皆保留於 data/ 之下,每個子資料夾對應一個原始來源 repo,例如: data/ ├── zhtw-literature-ots 每個子資料夾內保留: 原始 README 原始語料檔(json / txt 等) 來源資訊與授權說明 以利來源追溯與資料審核。 📦 資料格式 主要格式: JSON UTF-8 編碼文字檔 典型欄位可能包含: title:作品名稱 author:作者 content:文本內容 source:來源 repo (依各來源資料實際格式而定)… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/ima-corpus-zhtw.texttext-generation1K<n<10K0 likes12 downloads8mo agoHugging Face16IMA-Taiwan /taigi-literature-ttshsgated Dataset Summary The dataset contains 240 rows. These paragraphs are extracted from authorized novel written by Tiunn Tshing Siong張青松 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 240 (each representing a paragraph) Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ttshs.textn<1K0 likes11 downloads1y agoHugging Face17IMA-Taiwan /zhtw-literature-otsgated Dataset Summary The dataset contains 2,349 rows. These paragraphs are extracted from authorized novels written by Ou Tiong Siong胡長松 and contain multiple sentences in Traditional Chinese. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 2,349 (each representing a paragraph) Features: title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/zhtw-literature-ots.text1K<n<10K1 likes11 downloads1y agoHugging Face18IMA-Taiwan /taigi-literature-abtgated Dataset Summary The dataset contains 389 rows. These paragraphs are extracted from authorized novels written by Ang Bing-To洪明道 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 389 (each representing a paragraph) Features: title: Book… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-abt.textn<1K0 likes10 downloads2y agoHugging Face19IMA-Taiwan /taigi-literature-ngkhgated Dataset Summary The dataset contains 980 rows. These paragraphs are extracted from authorized paper written by Ngoo Ka Hun吳嘉芬 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 980 (each representing a paragraph) Features: title: Paper… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ngkh.textn<1K0 likes10 downloads1y agoHugging Face20IMA-Taiwan /taigi-literature-lgsgated Dataset Summary The dataset contains 190 rows. These paragraphs are extracted from authorized novels written by Lua Giok Si賴玉絲 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 190 (each representing a paragraph) Features: title: Book… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-lgs.textn<1K0 likes10 downloads1y agoHugging Face21IMA-Taiwan /taigi-literature-llbgated Dataset Summary The dataset contains 1,827 rows. These paragraphs are extracted from authorized novels written by Lîm lang-bín林央敏 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 1,827 (each representing a paragraph) Features: title:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-llb.text1K<n<10K0 likes10 downloads9mo agoHugging Face22IMA-Taiwan /taigi-literature-otsgated Dataset Summary The dataset contains 5,260 rows. These paragraphs are extracted from authorized novels written by Ou Tiong Siong胡長松 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 5,260 (each representing a paragraph) Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ots.text1K<n<10K6 likes9 downloads1y agoHugging Face23IMA-Taiwan /taigi-literature-kkhgated Dataset Summary The dataset contains 125 rows. These paragraphs are extracted from authorized novels written by Ko Ka-hui高嘉徽 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 125 (each representing a paragraph) Features: title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-kkh.textn<1K0 likes9 downloads1y agoHugging Face24IMA-Taiwan /taigi-literature-ljkgated Dataset Summary The dataset contains 579 rows. These paragraphs are extracted from authorized paper written by Lin Jui-Kun林瑞崐 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 579 (each representing a paragraph) Features: title: Paper… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ljk.textn<1K0 likes9 downloads2y agoHugging Face25IMA-Taiwan /taigi-literature-olbtgated Dataset Summary The dataset contains 617 rows. These paragraphs are extracted from authorized novels written by Ong Lo-Bit-To王羅蜜多 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 617 (each representing a paragraph) Features: title:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-olbt.textn<1K0 likes9 downloads2y agoHugging Face26IMA-Taiwan /taigi-literature-tksgated Dataset Summary The dataset contains 2,591 rows. These paragraphs are extracted from authorized novels written by Tan Kim-Sun陳金順 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 2,591 (each representing a paragraph) Features: title: Book… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-tks.text1K<n<10K0 likes8 downloads2y agoHugging Face27IMA-Taiwan /taigi-literature-tskgated Dataset Summary The dataset contains 84 rows. These paragraphs are extracted from authorized novels written by Tan Siu Ki陳秀枝 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 84 (each representing a paragraph) Features: title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-tsk.textn<1K0 likes8 downloads1y agoHugging Face28IMA-Taiwan /taigi-literature-ssltsgated Dataset Summary The dataset contains 443 rows. These paragraphs are extracted from authorized novels written by Sio Siann Ling Tsi小城綾子 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 443 (each representing a paragraph) Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-sslts.textn<1K0 likes8 downloads1y agoHugging Face29IMA-Taiwan /taigi-literature-achiakgated Dataset Summary The dataset contains 609 rows. These paragraphs are extracted from authorized prose written by Liau Tiunn Chiak廖張皭 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 609 (each representing a paragraph) Features: title:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-achiak.textn<1K0 likes8 downloads1y agoHugging Face30IMA-Taiwan /taigi-literature-khggated Dataset Summary The dataset contains 377 rows. These paragraphs are extracted from authorized novels written by Khng Guan康原 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 377 (each representing a paragraph) Features: title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-khg.textn<1K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.