CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JacobLinCool /taiwan-examsMachine-gradable exam benchmarks produced by any-to-bench. Each subset is one exam: the viewer table shows one row per answerable question (figures embedded); the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json (structured paper), answer_schema.json (strict JSON Schema an answer sheet must satisfy), grading.json (deterministic rules + judge rubrics), manifest.json (provenance), and assets/ (figures). Usage Benchmark any model against an exam: a2b download… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-exams.imagequestion-answering1K<n<10K0 likes2.7k downloads1mo agoHugging Face02aigrant /taiwan-ly-law-research Taiwan Legislator Yuan Law Research Data Overview The law research documents are issued irregularly from Taiwan Legislator Yuan. The purpose of those research are providing better understanding on social issues in aspect of laws. One may find documents rich with technical terms which could provided as training data. For comprehensive document list check out this link provided by Taiwan Legislator Yuan. There are currently missing document download links in 10th and 9th… See the full description on the dataset page: https://huggingface.co/datasets/aigrant/taiwan-ly-law-research.text1K<n<10K8 likes669 downloads10mo agoHugging Face03andynoodles /Taiwan-Financial Taiwan Financial Report OCR — block-level (zh-Hant) Block-level OCR pairs synthesised from Traditional-Chinese (zh-Hant) Taiwan listed-company financial reports. Each row is one cropped layout region with its block type and the vision-LLM transcription — including raw OTSL for tables. 194,314 rows from 138 financial reports (合併財報, IFRS consolidated, type AI1) 15 companies across major industries, 2021–2023, all four quarters Source: 公開資訊觀測站 / MOPS (doc.twse.com.tw) — public… See the full description on the dataset page: https://huggingface.co/datasets/andynoodles/Taiwan-Financial.documentimage-to-text100K<n<1M1 likes661 downloads3mo agoHugging Face04andynoodles /Taiwan-JudicialYuanPublication 司法周刊 Judicial Weekly OCR — block-level (zh-Hant, vertical text) Block-level OCR pairs synthesised from scanned issues of 司法周刊 (Judicial Weekly), the official weekly newspaper of Taiwan's Judicial Yuan (司法院). Each row is one cropped layout region with its block type and the vision-LLM transcription. 104,938 rows from 1,506 scanned pages (one PDF per 版-group, 期1–756) Complete scan era 1981–1995 (民國70–84), ~100 issues per year, every year covered Vertical Traditional Chinese (直書):… See the full description on the dataset page: https://huggingface.co/datasets/andynoodles/Taiwan-JudicialYuanPublication.documentimage-to-text10K<n<100K0 likes591 downloads3mo agoHugging Face05alix2t7 /Taiwan-black-bear-yolo 台灣黑熊偵測資料集 | Taiwan Black Bear Detection Dataset 🐻 用 AI 守護台灣國寶 🇹🇼 📖 資料集簡介 本資料集包含 18,112 張影像,專門用於訓練 台灣黑熊(學名:Ursus thibetanus formosanus)的物件偵測模型。所有標註均採用 YOLO 格式,可直接用於 YOLOv5、YOLOv8 等主流物件偵測框架。 🎯 核心特色 🌍 語言: 不適用(電腦視覺資料集) 📋 任務類型: 物件偵測(Object Detection) 🐾 應用領域: 野生動物保育、瀕危物種監測 💾 標註格式: YOLO v5/v8 相容格式 📊 類別數量: 1 類(台灣黑熊) ✨ 支援的應用場景 🔍 物件偵測: 在影像中偵測並定位台灣黑熊 📹 野外監測: 自動化野生動物追蹤與監控 🌲 保育研究: 支援瀕危物種保護工作 🚨 預警系統: 人熊衝突預防與警示 📊 資料集結構… See the full description on the dataset page: https://huggingface.co/datasets/alix2t7/Taiwan-black-bear-yolo.imageobject-detection10K<n<100K0 likes533 downloads7mo agoHugging Face06AoyamaAsake /QA_TaiwanEdoctortext100K<n<1M3 likes502 downloads2y agoHugging Face07sarahwei /Taiwanese-Minnan-Sutiau Taiwanese-Minnan-Sutiau Dataset The dataset consists of a curated collection of words that resemble tokens in Taiwanese Minnan (Taiwanese Hokkien), aimed at enhancing the recognition and processing of the language for various applications. Sourced from the Ministry of Education in Taiwan, this dataset serves as a valuable linguistic resource for researchers and developers engaged in language processing and recognition tasks. Dataset Features Source: Ministry of Education, Taiwan… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Sutiau.audioautomatic-speech-recognition10K<n<100K19 likes448 downloads2y agoHugging Face08andynoodles /Taiwan-LegislativeYuanPublication 立法院公報 Legislative Yuan Gazette OCR — block-level (zh-Hant) Block-level OCR pairs synthesised from scanned issues of the 立法院公報 (Legislative Yuan Official Gazette) of Taiwan, plus the 國民政府時期立法院公報 (Nationalist-government era, 1928–1943). Each row is one cropped layout region with its block type and the vision-LLM transcription — including raw OTSL for tables. 802,175 rows from 62,418 pages / 446 gazette volumes (冊) Letterpress & typewriter scans, spanning 卷40–98 (1951–2009) evenly… See the full description on the dataset page: https://huggingface.co/datasets/andynoodles/Taiwan-LegislativeYuanPublication.documentimage-to-text100K<n<1M0 likes440 downloads3mo agoHugging Face09adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-hokkien Taiwan-Tongues-ASR-CE-dataset-hokkien 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hokkien.audioautomatic-speech-recognition10K<n<100K7 likes422 downloads9mo agoHugging Face10hhhuang /TaiwanVQA TaiwanVQA: Benchmarking and Enhancing Cultural Understanding in Vision-Language Models Dataset Summary TaiwanVQA is a visual question answering (VQA) benchmark designed to evaluate the capability of vision-language models (VLMs) in recognizing and reasoning about culturally specific content related to Taiwan. This dataset contains 2,736 images captured by our team, paired with 5,472 manually designed questions that cover diverse topics from daily life in Taiwan… See the full description on the dataset page: https://huggingface.co/datasets/hhhuang/TaiwanVQA.imagevisual-question-answering10K<n<100K11 likes382 downloads10mo agoHugging Face11sarahwei /Taiwanese-Minnan-Example-Sentences Taiwanese Minnan Example Sentences The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems. Dataset Features Source: Ministry of Education, Taiwan (Sutian Resource Center) Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.audioautomatic-speech-recognition10K<n<100K12 likes301 downloads2y agoHugging Face12EZCon /taiwan-license-plate-recognitionimageimage-segmentation1K<n<10K2 likes274 downloads11mo agoHugging Face13skyhong2002 /taiwan-professional-exams-115-2Machine-gradable exam benchmarks produced by any-to-bench. Each subset is one exam: the viewer table shows one row per answerable question (figures embedded); the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json (structured paper), answer_schema.json (strict JSON Schema an answer sheet must satisfy), grading.json (deterministic rules + judge rubrics), manifest.json (provenance), and assets/ (figures). Usage Benchmark any model against an exam: a2b download… See the full description on the dataset page: https://huggingface.co/datasets/skyhong2002/taiwan-professional-exams-115-2.imagequestion-answering1K<n<10K0 likes233 downloads1mo agoHugging Face14yhzkiki /taiwan-dtm-2025-terrarium-z13 2025 年版全臺灣 20 m DTM — Terrarium z13 這個 Dataset 將內政部公開的 2025 年版全臺灣 20 公尺網格數值地形模型(DTM)轉為 ShadeMap 可直接讀取的 Terrarium RGB XYZ tiles。 官方資料來源:https://data.gov.tw/dataset/176927 原始資料授權:政府資料開放授權條款-第 1 版。本 repository 為衍生格式,請保留官方來源與授權資訊。 內容 terrain/13/{x}/{y}.png:256×256 RGB PNG,XYZ / Web Mercator tile addressing。 tile-index.csv:每張 tile 的區域、有效像素比例、來源高程範圍與 Terrarium 量化誤差。 build-summary.json:建置摘要。 source-manifest.json:原始 ZIP/TIFF SHA-256、解析度、範圍與 CRS 決策。… See the full description on the dataset page: https://huggingface.co/datasets/yhzkiki/taiwan-dtm-2025-terrarium-z13.imagen<1K0 likes233 downloads11d agoHugging Face15yentinglin /TaiwanChat Performance Citation If you find Taiwan LLM is useful in your work, please cite it with: @misc{lin2023taiwan, title={Taiwan LLM: Bridging the Linguistic Divide with a Culturally Aligned Language Model}, author={Yen-Ting Lin and Yun-Nung Chen}, year={2023}, eprint={2311.17487}, archivePrefix={arXiv}, primaryClass={cs.CL} } texttext-generation100K<n<1M69 likes229 downloads2y agoHugging Face16yhzkiki /taiwan-shade-routing-data Taiwan Shade Routing Runtime Tiles — dev26 Browser-runtime dataset for the dev26 nationwide route/shade prototype. The website lazy-loads only the HDTL evidence tiles and HGR1 graph tiles around the active A→B route; it does not download national archives into the browser. Finalized build Combined tiles: 1625 Official evidence tiles: 1157 Routing graph tiles: 1464 Official accepted features: 94140 Sidewalk features: 92192 Bikeway features: 1948 Pedestrian… See the full description on the dataset page: https://huggingface.co/datasets/yhzkiki/taiwan-shade-routing-data.geospatial0 likes225 downloads2d agoHugging Face17JacobLinCool /taiwan-exams-resultsBenchmark results produced by any-to-bench. One subset here is one taker configuration — a single model at a single reasoning effort — sat against the exams in another dataset repo. Every row names the exam repo and subset it was earned against, so results from several corpora, and from several people, can live side by side. results-index.json — the catalog: one headline row per configuration results-<entry>/entry.json — that configuration's per-paper scores results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-exams-results.tabularquestion-answering10K<n<100K0 likes220 downloads17d agoHugging Face18adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-zhtw Taiwan-Tongues-ASR-CE-dataset-zhtw 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。 🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-zhtw.audioautomatic-speech-recognition100K<n<1M2 likes208 downloads9mo agoHugging Face19hcy-43 /Taiwan-Excellent-with-Webcrawltext1M<n<10M0 likes204 downloads1y agoHugging Face20aqweteddy /Taiwan-VQA-Train-Messagesimage1K<n<10K0 likes201 downloads1y agoHugging Face21EZCon /taiwan-license-plate-detectionimage1K<n<10K0 likes180 downloads1y agoHugging Face22lianghsun /QA_TaiwanEdoctor Dataset Card for QA_TaiwanEdoctor QA_TaiwanEdoctor 是一個以台灣線上醫療諮詢平台為來源之繁體中文醫療問答資料集,合計 178,126 筆,時間跨度 2000–2024 年。每筆包含使用者之健康提問(ask)與醫師之回覆(ans),並附上問題標題、提問日期、瀏覽次數與(可選之)評分,可作為繁中醫療 LLM 之 CPT / SFT 訓練素材。 Dataset Details Dataset Description 繁體中文醫療問答語料長期稀缺,使得繁中 LLM 在健康諮詢場景之表現受限。本資料集整理自台灣公開線上醫療問答平台,內容涵蓋內科、外科、眼科、皮膚科、兒科、婦產科、精神科、牙科等各科別,由執業醫師回覆使用者之健康問題。 資料以 312 個 JSON 檔案儲存,每檔包含數百至上千筆 QA。每筆欄位: title:問題標題(含編號); ask:使用者之健康提問(自由文字); ans:醫師之回覆; qtime:提問日期(格式 YYYY/MM/DD);… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/QA_TaiwanEdoctor.textquestion-answering100K<n<1M0 likes175 downloads5mo agoHugging Face23twinkle-ai /Llama-3-Taiwan-70B-Instruct-eval-logs-and-scorestabular100K<n<1M0 likes172 downloads7mo agoHugging Face24twinkle-ai /Llama-3.1-Taiwan-8B-Instruct-eval-logs-and-scorestabular100K<n<1M0 likes168 downloads7mo agoHugging Face25JacobLinCool /taiwan-national-exams-resultsBenchmark results produced by any-to-bench. One subset here is one taker configuration — a single model at a single reasoning effort — sat against the exams in another dataset repo. Every row names the exam repo and subset it was earned against, so results from several corpora, and from several people, can live side by side. results-index.json — the catalog: one headline row per configuration results-<entry>/entry.json — that configuration's per-paper scores results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-national-exams-results.tabularquestion-answering1K<n<10K0 likes162 downloads1mo agoHugging Face26andynoodles /Taiwan-UrbanPlan 台灣都市計畫書 OCR — block-level (zh-Hant) Block-level OCR pairs synthesised from scanned Taiwanese urban-plan books (都市計畫書) published by county governments. Each row is one cropped layout region with its block type and the vision-LLM transcription — including raw OTSL for tables. 119,735 rows from 15,954 pages / 318 plan books (計畫書 + some 計畫圖) 36 urban-planning districts (都計區) in 屏東縣 (Pingtung) and 嘉義市 (Chiayi City), plan dates spanning 民國40年代–110年代 (1950s–2020s) Mostly… See the full description on the dataset page: https://huggingface.co/datasets/andynoodles/Taiwan-UrbanPlan.documentimage-to-text100K<n<1M1 likes157 downloads3mo agoHugging Face27justicedao /ipfs_taiwan_laws_ir Taiwan legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_taiwan_laws (revision 4b2305d5d9ae9ae8796be3d36ceab42d53dd6f78) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Taiwan prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_taiwan_laws_ir.tabulartext-retrieval1M<n<10M0 likes130 downloads2d agoHugging Face28adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-hakka Taiwan-Tongues-ASR-CE-dataset-hakka 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。 🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hakka.audioautomatic-speech-recognition1K<n<10K2 likes125 downloads9mo agoHugging Face29benchang1110 /Vision-Taiwan-595kimage100K<n<1M0 likes124 downloads2y agoHugging Face30zkdeng /taiwanSpiders Dataset Card for "taiwanSpiders" More Information needed image10K<n<100K0 likes120 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.