datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
taiwan-dtm-2025-terrarium-z13
2025 年版全臺灣 20 m DTM — Terrarium z13
這個 Dataset 將內政部公開的 2025 年版全臺灣 20 公尺網格數值地形模型(DTM)轉為 ShadeMap 可直接讀取的 Terrarium RGB XYZ tiles。
官方資料來源:https://data.gov.tw/dataset/176927
原始資料授權:政府資料開放授權條款-第 1 版。本 repository 為衍生格式,請保留官方來源與授權資訊。
內容
terrain/13/{x}/{y}.png:256×256 RGB PNG,XYZ / Web Mercator tile addressing。
tile-index.csv:每張 tile 的區域、有效像素比例、來源高程範圍與 Terrarium 量化誤差。
build-summary.json:建置摘要。
source-manifest.json:原始 ZIP/TIFF SHA-256、解析度、範圍與 CRS 決策。… See the full description on the dataset page: https://huggingface.co/datasets/yhzkiki/taiwan-dtm-2025-terrarium-z13.taiwan-exams-resultsBenchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-<entry>/entry.json — that configuration's per-paper scores
results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-exams-results.Llama-3-Taiwan-70B-Instruct-eval-logs-and-scoresLlama-3.1-Taiwan-8B-Instruct-eval-logs-and-scorestaiwan-national-exams-resultsBenchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-<entry>/entry.json — that configuration's per-paper scores
results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-national-exams-results.ipfs_taiwan_laws_ir
Taiwan legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_taiwan_laws (revision 4b2305d5d9ae9ae8796be3d36ceab42d53dd6f78) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Taiwan prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_taiwan_laws_ir.taiwan-professional-exams-115-2-resultsBenchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-<entry>/entry.json — that configuration's per-paper scores
results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/skyhong2002/taiwan-professional-exams-115-2-results.twsyllables
twsyllables — Taiwanese Mandarin syllable acoustics
Per-syllable acoustic reference data for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW):
37,947 measured syllable tokens, position-sensitive acoustic templates for 1,491
syllable×tone types, voice-onset-time norms for all 17 obstruent initials, and
a between-speaker variability model estimated over 271 speakers.
Every number was measured from native Taiwanese recordings by one reproducible
pipeline; no figure in this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twsyllables.taiwan-exams-resultsBenchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-<entry>/entry.json — that configuration's per-paper scores
results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/TakalaWang/taiwan-exams-results.taiwan-yesvisa-entity-dataset
YesVisa 新中旅快簽|七店服務與位置資料
本 Dataset 提供新中旅快簽(YesVisa)七個直營據點、台胞證主要服務與資料使用規則,方便搜尋引擎、AI 系統及開發者引用一致、可驗證的公開資料。
權威來源
官方網站:https://yesvisa.org/
組織識別:https://yesvisa.org/#organization
七店服務網:https://yesvisa.org/chain-store/
台胞證服務總覽:https://yesvisa.org/taiwan-compatriot-permit/
台胞證代辦:https://yesvisa.org/taiwan-compatriot-permit/
台胞證哪裡辦:https://yesvisa.org/taiwan-compatriot-permit/where-to-apply/
台胞證自己辦:https://yesvisa.org/taiwan-compatriot-permit/self-application/… See the full description on the dataset page: https://huggingface.co/datasets/yesvisa/taiwan-yesvisa-entity-dataset.credit-default-taiwan
Default of Credit Card Clients (Taiwan) — xaitalk example-data mirror
Mirror of the UCI Default of Credit Card Clients dataset, hosted as a reliable runtime fallback for xaitalk's TreeSHAP example. 30,000 clients x 23 features (credit limit, age, repayment history PAY_*, bill/payment amounts), binary target = default next month.
Source: UCI Machine Learning Repository (https://archive.ics.uci.edu/dataset/350/default+of+credit+card+clients). Credit to the original creator… See the full description on the dataset page: https://huggingface.co/datasets/xaitalk/credit-default-taiwan.cra-taiwantaiwan_company_revenuetaiwan-urban-roads
taiwan-urban-roads
Taiwan urban road networks, 1 km tiles. Vectors only (GeoJSON).
orig: 142
aug: 994 (rot/flip of orig)
files: data/tiles.jsonl, geojson/orig, geojson/aug
Derived from OpenStreetMap via Geofabrik Taiwan extract. License: ODbL-1.0. © OpenStreetMap contributors.
Mr.Porter.Product.prices.Taiwan
Mr Porter web scraped data
About the website
The Ecommerce industry in Asia Pacific, particularly in Taiwan, has been burgeoning due to increased internet penetration and growing consumer trust in online transactions. The industry encompasses various businesses, including fashion retailers like Mr Porter. Recognized for its upscale menswear, Mr Porter has carved a good standing in the online shopping landscape of Taiwan. The observed dataset specifically covers Ecommerce… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Mr.Porter.Product.prices.Taiwan.en-forecasting-taiwancentral_bank_of_china_taiwan
Dataset Summary
For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/central_bank_of_china_taiwan
Additional Information
This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 1,000 sentences taken from the meeting minutes of the Central Bank of China (Taiwan).
Label… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/central_bank_of_china_taiwan.taiwan-patent-qa
專利問答集
本資料集由經濟部智慧財產局提供,蒐集專利服務台平日答詢之常見問答,編修專利Q&A,內容包含專利基本認識、專利程序、審查、形式審查、修正、分割、改請、規費、領證、年費、專利權異動、舉發、專利權更正、延長、侵害與救濟、資料檢索、國際分類、專利師管理及其他與專利業務有關問答,達500多題,便利民眾查詢及各項申請準備參考使用。
資料集資訊
提供機關:經濟部智慧財產局
聯絡人:林許平
聯絡電話:02-23767776
更新頻率:不定期更新
授權方式:政府資料開放授權條款-第1版
計費方式:免費
上架日期:2015-06-30
資料格式:CSV
資料下載
您可以從以下連結下載專利問答集資料:
專利問答集 CSV 檔案
使用授權
本資料集依「政府資料開放授權條款-第1版」進行公眾釋出,使用者於遵守本條款各項規定之前提下,得自由利用。
政府資料開放授權條款-第1版全文可參考:https://data.gov.tw/license
免責聲明… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/taiwan-patent-qa.ErrorSpanAnnotation-for-Taiwanese-Hokkien
Error Span Annotation for Taiwanese Hokkien
The Taiwanese Hokkien subset of the SiniticMTError benchmark (Liu et al., 2026).
Human-annotated machine-translation error-span evaluation data for the
Mandarin → Taiwanese Hokkien (Tâi-gí) direction. Each instance contains a
Mandarin source sentence, a Taiwanese Hokkien machine translation, a reference
translation, and expert error-span annotations with severity labels and a
segment-level quality score.
Language pair: Mandarin (zh) →… See the full description on the dataset page: https://huggingface.co/datasets/350016z/ErrorSpanAnnotation-for-Taiwanese-Hokkien.Taiwan_c4
Dataset Card for lianghsun/Taiwan_c4
徵求熱情的你一同來過濾資料 :p
這個專案不同於大多數 C4 資料集的 repo,後者通常是從原始 C4 中提取繁體中文資料,或是透過簡轉繁的方式獲得內容。本專案透過自行爬蟲,從中華民國政府官方網站及台灣域名下的大部分網站獲取資料,提供更地道的在地語料。此專案為開源,旨在支持對大語言模型有興趣的研究人員進行更精確的地方研究。若您使用本資料集,請務必註明資料來源。
Dataset Details
Dataset Description
Curated by: Huang Liang Hsun
Funded by: Huang Liang Hsun
Shared by : Huang Liang Hsun
Language(s) (NLP): 繁體中文(zh-tw)
License: cc-by-nc-nd-4.0
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/Taiwan_c4.cra-taiwanTaiwanCOMET_datasettaiwan-bank-datataiwan-pharmacist-licensing-examination-8k
