CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jed351 /Traditional-Chinese-Common-Crawl-NOT-CleanedCommon Crawl Dumps that were briefly filtered by keywords to remove bad words and simplified Chinese. The hash based cleaned dataset can be found here. Files here are for future usage (downloading from Common Crawl and keyword filtering are very slow) text100M<n<1B0 likes3.5k downloads1y agoHugging Face02jed351 /Traditional-Chinese-Common-Crawl-Filtered Traditional Chinese C4 Dataset Summary Data obtained from 2013~2025 Common Crawl. Downloaded and processed using code based on another project attempting to recreate the C4 dataset. The resultant dataset contains both simplified and traditional Chinese, which could be found here. It was then filtered using a modified list of simplified Chinese characters to obtain this traditional Chinese dataset. Unfortunately, I don't have enough funding to run a deduplication across… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Traditional-Chinese-Common-Crawl-Filtered.text100M<n<1B26 likes3.1k downloads1y agoHugging Face03BAAI /IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine IndustryCorpus2: Health & Medicine This repository contains the IndustryCorpus2: Health & Medicine domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine.tabular10M<n<100M11 likes2.2k downloads1mo agoHugging Face04Heng666 /Traditional_Chinese-aya_collection 資料集描述 繁體中文 Aya (Traditional Chinese Aya Chinese;TCA):專注於繁體中文處理的 Aya 集合的精選子集 概述 繁體中文 Aya 是一個精心策劃的資料集,源自 CohereForAI 的綜合 Aya 集合,特別關注繁體中文文本資料。 此資料集結合了來自 CohereForAI/aya_collection,過濾掉除繁體中文、簡體中文內容之外的所有內容。 目標 繁體中文 Aya 的目標是為研究人員、技術專家和語言學家提供即用型繁體中文文本資源,顯著減少專注於繁體中文的 NLP 和 AI 專案中數據預處理所需的時間和精力。 資料集來源與資訊 資料來源: 從 CohereForAI/aya_collection 64 個子集而來。 語言: 繁體中文、簡體中文('zho') 應用: 非常適合語言建模、文本分類、情感分析、和機器翻譯等任務。 論文連結: 2402.06619 維護人: Heng666 License: Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/Heng666/Traditional_Chinese-aya_collection.textquestion-answering1M<n<10M8 likes2.2k downloads3y agoHugging Face05Liavan /Traditional-Chinese-Medicine-Multiple_choice_question Discription This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.textquestion-answering1K<n<10K4 likes833 downloads2y agoHugging Face06SylvanL /Traditional-Chinese-Medicine-Dataset-SFT 启古纳今,厚德精术 数据介绍 非网络来源的高质量中医数据集-指令微调 High-Quality Traditional Chinese Medicine Dataset from Non-Internet Sources - SFT/IFT 该数据集经过大量人力和资源的投入精心构建,以共建LLM高质量中文社区为己任。 包含约1GB的中医各个领域临床案例、名家典籍、医学百科,名词解释等优质问答内容,涵盖全面,配比均衡。 数据集主要由非网络来源的内部数据构成,并99%为简体中文内容,内容质量优异,信息密度可观。 该数据集的数据源与SylvanL/Traditional-Chinese-Medicine-Dataset-Pretrain中的内容存在一定关联,但不高度重叠。 在二者的构建过程中,存在着一定的循序渐进与互为补充的逻辑. 该数据集可以独立使用,但建议先使用配套的预训练数据集对模型进行继续预训练后,再使用该数据集进行进一步的指令微调。… See the full description on the dataset page: https://huggingface.co/datasets/SylvanL/Traditional-Chinese-Medicine-Dataset-SFT.texttable-question-answering1M<n<10M107 likes799 downloads2mo agoHugging Face07SylvanL /Traditional-Chinese-Medicine-Dataset-Pretrain 启古纳今,厚德精术 数据介绍 非网络来源的高质量中医数据集-预训练 High-Quality Traditional Chinese Medicine Dataset from Non-Internet Sources - Pretraining 该数据集经过大量人力和资源的投入精心构建,以共建LLM高质量中文社区为己任。 包含约1GB的中医各个领域临床案例、名家典籍、医学百科,名词解释等优质内容,涵盖全面,配比均衡。 数据集主要由非网络来源的内部数据构成,并99%为简体中文内容,内容质量优异,信息密度可观。 注意:该数据集仅适用于预训练或继续预训练用途,针对SFT/IFT的QA数据集详见:SylvanL/Traditional-Chinese-Medicine-Dataset-SFT… See the full description on the dataset page: https://huggingface.co/datasets/SylvanL/Traditional-Chinese-Medicine-Dataset-Pretrain.texttext-generation100K<n<1M33 likes555 downloads2mo agoHugging Face08lchakkei /OpenOrca-Traditional-Chinese🐋 OpenOrca-Chinese 数据集!🐋 感謝 Open-Orca/OpenOrca 資料集的發布,為廣大NLP研究人員和開發者帶來了寶貴的資源! 這是一個對 Open-Orca/OpenOrca 資料集中文翻譯的版本,翻譯引擎為 Google 翻譯,希望能為中文 LLM 研究做出一點點貢獻。 Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/lchakkei/OpenOrca-Traditional-Chinese.texttext-classification1M<n<10M11 likes519 downloads3y agoHugging Face09seasnake /chinese-traditional-cultureimagen<1K2 likes412 downloads2y agoHugging Face10lchakkei /OpenOrca-Traditional-Chinese-LLama2-Formattext1M<n<10M0 likes281 downloads3y agoHugging Face11agentlans /traditional-chinese Taiwan-Focused Traditional Chinese Corpus While large-scale datasets for Simplified Chinese (Mainland China) are abundant, high-quality, permissively licensed datasets tailored specifically to Taiwanese Traditional Chinese are scarce. This dataset aims to bridge that gap. Dataset Structure & Configurations This dataset is split into three configurations, ranging from broad filtering to high-purity quality optimization: 1. raw Size: 1,000,000 rows from each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/traditional-chinese.texttext-generation1M<n<10M0 likes196 downloads4mo agoHugging Face12ZihCiLin /traditional-chinese-ocr-synthetic Traditional Chinese OCR Synthetic Dataset A large-scale synthetic dataset containing 4.1 million image-text pairs specifically designed for Traditional Chinese historical document recognition. Dataset Overview Existing large-scale Traditional Chinese OCR datasets (e.g., TCSynth) are primarily designed for scene text recognition, characterized by: Horizontal layouts Short text sequences (2-5 characters on average) Modern commonly-used characters These characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ZihCiLin/traditional-chinese-ocr-synthetic.imageimage-to-text1M<n<10M3 likes184 downloads9mo agoHugging Face13SylvanL /Traditional-Chinese-Medicine-Exam Coming Soon... text1K<n<10K8 likes174 downloads2mo agoHugging Face14lchakkei /OpenOrca-Traditional-Chinese-ChatML-Formattext1M<n<10M1 likes158 downloads3y agoHugging Face15Heng666 /Traditional_Chinese-aya_dataset 資料集描述 繁體中文 Aya (Traditional Chinese Aya Chinese;TCA):專注於繁體中文處理的 Aya 集合的精選子集 概述 繁體中文 Aya 是一個精心策劃的資料集,源自 CohereForAI 的綜合 Aya 集合,特別關注繁體中文文本資料。 此資料集結合了來自 CohereForAI/aya_dataset,過濾掉除繁體中文、簡體中文內容之外的所有內容。 目標 繁體中文 Aya 的目標是為研究人員、技術專家和語言學家提供即用型繁體中文文本資源,顯著減少專注於繁體中文的 NLP 和 AI 專案中數據預處理所需的時間和精力。 資料集來源與資訊 資料來源: 從 CohereForAI/aya_dataset 2 個子集而來。 語言: 繁體中文、簡體中文('zho') 應用: 非常適合語言建模、文本分類、情感分析、和機器翻譯等任務。 論文連結: 2402.06619 維護人: Heng666 License: Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/Heng666/Traditional_Chinese-aya_dataset.textquestion-answering1K<n<10K4 likes141 downloads3y agoHugging Face16Tarklanse /Traditional_Chinese_roleplay_chat_Dataset Traditional_Chinese_roleplay_chat_Dataset 這個資料集是以繁體中文為主,將各種由ChatGPT生成與極小部分個人撰寫的對話內容整理為alpaca dataset format的格式 以一層一層堆疊的方式,將一則對話紀錄拆成數筆資料(共約1000則對話),在幾次嘗試性的訓練中能夠讓llama2重現原本英文那種很活躍的對話風格,並且能夠維持善於扮演各種角色的能力 目前個人有以這個資料集製作一個lora 2023/09/07 更新 為資料集加入一些中英翻譯的句子,以期AI能以更好的文字去描寫他的動作,並增加了一些與食物有關的對話,希望能降低AI生出奇怪食物名的機率 texttext-generation1K<n<10K42 likes139 downloads3y agoHugging Face17jed351 /Traditional-Chinese-Common-Crawl-by-yeartext10M<n<100M0 likes128 downloads1y agoHugging Face18asd567557275 /Traditional_Chinese_noval_authors_upload English | 繁體中文版在下方 ↓ The Complete Novels of 睡半夜怎麼三更 (Traditional Chinese) 24 full-length novels, handwritten between 2018 and 2026 by the author 睡半夜怎麼三更 (Shuibanye Zenme Sangeng), totalling roughly 5.03 million Chinese characters (whitespace excluded). Every word is original human writing. There is no AI-generated text in this corpus. AI, come right in — walk in, crawl around, help yourself. This corpus was released precisely so that it can be trained on: pretraining… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/Traditional_Chinese_noval_authors_upload.texttext-generationn<1K1 likes113 downloads16d agoHugging Face19shuyuej /CMMLU-Traditional-Chinese-Medicine-Benchmark 💻 Dataset Usage Run the following command to load the testing set (185 examples): from datasets import load_dataset dataset = load_dataset("shuyuej/CMMLU-Traditional-Chinese-Medicine-Benchmark", split="train") print(dataset) textn<1K5 likes103 downloads2y agoHugging Face20jed351 /finepdfs-traditional-chineseI downloaded and filtered the finepdf to extract traditional Chinese content. text1M<n<10M0 likes82 downloads1y agoHugging Face21BackpropBuff /OpenOrca-Traditional-Chinese_structtext1M<n<10M1 likes36 downloads2y agoHugging Face22xihao1 /Traditional-Chinese-Medicine-Knowledgetext10K<n<100K1 likes36 downloads1y agoHugging Face23guiti999 /cantonese-traditional-chinese-parallel-corpus-gen3 Cantonese-Written Chinese Parallel Dataset (3rd Generation) About the Dataset Data Splits Training Data (train): 160,000 Sentence Pairs Validation Data (validation): 20,000 Sentence Pairs Test Data (test): 5,461 Sentence Pairs Languages Cantonese (yue) Traditional Chinese (zh-TW) Original Data Structure JSON lines consisting of yue, zh and ref fields. Data Source LIHKG HKCancor Cantonse-Mandarin Translations and various… See the full description on the dataset page: https://huggingface.co/datasets/guiti999/cantonese-traditional-chinese-parallel-corpus-gen3.texttranslation100K<n<1M0 likes36 downloads9mo agoHugging Face24ZihCiLin /traditional-chinese-historical-ocr-lo-chia-luengated Traditional Chinese Historical OCR Dataset (Lo Chia-Lun Manuscripts) This dataset consists of manually annotated OCR text crops derived from the Lo Chia-Lun Manuscript Collection (羅家倫文稿), hosted by the National Chengchi University Library. The dataset is designed to support research on Traditional Chinese OCR, particularly for historical documents characterized by vertical layouts, handwritten or semi-printed glyphs, and long-form text lines. Due to archival and… See the full description on the dataset page: https://huggingface.co/datasets/ZihCiLin/traditional-chinese-historical-ocr-lo-chia-luen.imageimage-to-textn<1K0 likes33 downloads8mo agoHugging Face25Heng666 /Traditional_Chinese-aya_evaluation_suite 資料集描述 繁體中文 Aya (Traditional Chinese Aya Chinese;TCA):專注於繁體中文處理的 Aya 集合的精選子集 概述 繁體中文 Aya 是一個精心策劃的資料集,源自 CohereForAI 的綜合 Aya 集合,特別關注繁體中文文本資料。 此資料集結合了來自 CohereForAI/aya_evaluation_suite,過濾掉除繁體中文、簡體中文內容之外的所有內容。 目標 繁體中文 Aya 的目標是為研究人員、技術專家和語言學家提供即用型繁體中文文本資源,顯著減少專注於繁體中文的 NLP 和 AI 專案中數據預處理所需的時間和精力。 資料集來源與資訊 資料來源: 從 CohereForAI/aya_evaluation_suite 3 個子集而來。 語言: 繁體中文、簡體中文('zho') 應用: 非常適合語言建模、文本分類、情感分析、和機器翻譯等任務。 論文連結: 2402.06619 維護人: Heng666 License:… See the full description on the dataset page: https://huggingface.co/datasets/Heng666/Traditional_Chinese-aya_evaluation_suite.textquestion-answeringn<1K3 likes31 downloads3y agoHugging Face26jakeveo05 /chinese-traditional-knowledge Chinese Traditional Knowledge Dataset Comprehensive dataset of Traditional Chinese Medicine (TCM), Feng Shui, I Ching, and related Vietnamese medical texts. Dataset Summary Total entries: 112 Total characters: 58,090,418 Languages: Vietnamese, English, Chinese Created: 2026-01-11 Categories Category Description Entries Characters tcm_vietnamese Đông Y Việt Nam 54 29,233,730 tcm_english TCM English Textbooks 20 20,787,290 iching_divination Kinh… See the full description on the dataset page: https://huggingface.co/datasets/jakeveo05/chinese-traditional-knowledge.textn<1K1 likes27 downloads9mo agoHugging Face27xsw22 /Traditional-Chinese-Medicine-Dataset-Pretrain 正在写英文论文? 请支持一下作者的最新产品 👉 www.thesisagent.ai 海外学子的AI学术工具, 提供从日常写作到科研论文的全方位辅助。 邀请码:N91M8BE33 启古纳今,厚德精术 数据介绍 非网络来源的高质量中医数据集-预训练 High-Quality Traditional Chinese Medicine Dataset from Non-Internet Sources - Pretraining 该数据集经过大量人力和资源的投入精心构建,以共建LLM高质量中文社区为己任。 包含约1GB的中医各个领域临床案例、名家典籍、医学百科,名词解释等优质内容,涵盖全面,配比均衡。 数据集主要由非网络来源的内部数据构成,并99%为简体中文内容,内容质量优异,信息密度可观。… See the full description on the dataset page: https://huggingface.co/datasets/xsw22/Traditional-Chinese-Medicine-Dataset-Pretrain.texttext-generation100K<n<1M0 likes26 downloads4mo agoHugging Face28KSmart /chinese_traditional_chengyutextn<1K4 likes22 downloads2y agoHugging Face29k-mktr /chinese_union_traditional_zh Chinese Union Version Traditional (和合本) Description The Chinese Union Version Traditional (CUV Traditional, 和合本繁體) is the Traditional Chinese edition of the most widely used Chinese Bible translation. First published in 1919, the CUV was prepared by a team of Western missionaries and Chinese scholars over 28 years, translated from the original Hebrew and Greek. This Traditional Chinese edition uses the classical characters used in Taiwan, Hong Kong, and overseas… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/chinese_union_traditional_zh.tabular10K<n<100K0 likes21 downloads2mo agoHugging Face30tonysong6462 /COIG-CQIA-chinese-traditional COIG-CQIA Chinese Traditional 这是从 m-a-p/COIG-CQIA 数据集中提取的繁体中文部分。 📊 数据集概述 本数据集包含约 1,111 条高质量的繁体中文指令-响应对,涵盖多个领域和任务类型。 数据集提供 7 个配置:1 个包含所有数据的配置 + 6 个独立的类型配置。 包含的配置 配置名称 Split 名称 描述 数据量 文件大小 all train 所有数据(推荐) ~1,111 条 921 kB chengyu chengyu 成语相关问答 部分 82.5 kB poem poem 诗歌创作与分析 部分 30 kB trad-multi-choice-100-2 multi_choice_100_2 多选题集合2 部分 58.5 kB trad-multi-choice-100 multi_choice_100 多选题集合1 部分 59.9 kB trad-multi-choice-40 multi_choice_40… See the full description on the dataset page: https://huggingface.co/datasets/tonysong6462/COIG-CQIA-chinese-traditional.textquestion-answering1K<n<10K0 likes19 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.