datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Traditional-Chinese-Common-Crawl-NOT-CleanedCommon Crawl Dumps that were briefly filtered by keywords to remove bad words and simplified Chinese.
The hash based cleaned dataset can be found here.
Files here are for future usage (downloading from Common Crawl and keyword filtering are very slow)
Traditional-Chinese-Common-Crawl-Filtered
Traditional Chinese C4
Dataset Summary
Data obtained from 2013~2025 Common Crawl.
Downloaded and processed using code based on another project attempting to recreate the C4 dataset.
The resultant dataset contains both simplified and traditional Chinese, which could be found here.
It was then filtered using a modified list of simplified Chinese characters to obtain this traditional Chinese dataset.
Unfortunately, I don't have enough funding to run a deduplication across… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Traditional-Chinese-Common-Crawl-Filtered.IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine
IndustryCorpus2: Health & Medicine
This repository contains the IndustryCorpus2: Health & Medicine domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine.Traditional_Chinese-aya_collection
資料集描述
繁體中文 Aya (Traditional Chinese Aya Chinese;TCA):專注於繁體中文處理的 Aya 集合的精選子集
概述
繁體中文 Aya 是一個精心策劃的資料集,源自 CohereForAI 的綜合 Aya 集合,特別關注繁體中文文本資料。
此資料集結合了來自 CohereForAI/aya_collection,過濾掉除繁體中文、簡體中文內容之外的所有內容。
目標
繁體中文 Aya 的目標是為研究人員、技術專家和語言學家提供即用型繁體中文文本資源,顯著減少專注於繁體中文的 NLP 和 AI 專案中數據預處理所需的時間和精力。
資料集來源與資訊
資料來源: 從 CohereForAI/aya_collection 64 個子集而來。
語言: 繁體中文、簡體中文('zho')
應用: 非常適合語言建模、文本分類、情感分析、和機器翻譯等任務。
論文連結: 2402.06619
維護人: Heng666
License: Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/Heng666/Traditional_Chinese-aya_collection.Traditional-Chinese-Medicine-Multiple_choice_question
Discription
This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.Traditional-Chinese-Medicine-Dataset-SFT
启古纳今,厚德精术
数据介绍
非网络来源的高质量中医数据集-指令微调
High-Quality Traditional Chinese Medicine Dataset from Non-Internet Sources - SFT/IFT
该数据集经过大量人力和资源的投入精心构建,以共建LLM高质量中文社区为己任。
包含约1GB的中医各个领域临床案例、名家典籍、医学百科,名词解释等优质问答内容,涵盖全面,配比均衡。
数据集主要由非网络来源的内部数据构成,并99%为简体中文内容,内容质量优异,信息密度可观。
该数据集的数据源与SylvanL/Traditional-Chinese-Medicine-Dataset-Pretrain中的内容存在一定关联,但不高度重叠。
在二者的构建过程中,存在着一定的循序渐进与互为补充的逻辑.
该数据集可以独立使用,但建议先使用配套的预训练数据集对模型进行继续预训练后,再使用该数据集进行进一步的指令微调。… See the full description on the dataset page: https://huggingface.co/datasets/SylvanL/Traditional-Chinese-Medicine-Dataset-SFT.Traditional-Chinese-Medicine-Dataset-Pretrain
启古纳今,厚德精术
数据介绍
非网络来源的高质量中医数据集-预训练
High-Quality Traditional Chinese Medicine Dataset from Non-Internet Sources - Pretraining
该数据集经过大量人力和资源的投入精心构建,以共建LLM高质量中文社区为己任。
包含约1GB的中医各个领域临床案例、名家典籍、医学百科,名词解释等优质内容,涵盖全面,配比均衡。
数据集主要由非网络来源的内部数据构成,并99%为简体中文内容,内容质量优异,信息密度可观。
注意:该数据集仅适用于预训练或继续预训练用途,针对SFT/IFT的QA数据集详见:SylvanL/Traditional-Chinese-Medicine-Dataset-SFT… See the full description on the dataset page: https://huggingface.co/datasets/SylvanL/Traditional-Chinese-Medicine-Dataset-Pretrain.OpenOrca-Traditional-Chinese🐋 OpenOrca-Chinese 数据集!🐋
感謝 Open-Orca/OpenOrca 資料集的發布,為廣大NLP研究人員和開發者帶來了寶貴的資源!
這是一個對 Open-Orca/OpenOrca 資料集中文翻譯的版本,翻譯引擎為 Google 翻譯,希望能為中文 LLM 研究做出一點點貢獻。
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/lchakkei/OpenOrca-Traditional-Chinese.chinese-traditional-cultureOpenOrca-Traditional-Chinese-LLama2-Formattraditional-chinese
Taiwan-Focused Traditional Chinese Corpus
While large-scale datasets for Simplified Chinese (Mainland China) are abundant, high-quality, permissively licensed datasets tailored specifically to Taiwanese Traditional Chinese are scarce.
This dataset aims to bridge that gap.
Dataset Structure & Configurations
This dataset is split into three configurations, ranging from broad filtering to high-purity quality optimization:
1. raw
Size: 1,000,000 rows from each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/traditional-chinese.traditional-chinese-ocr-synthetic
Traditional Chinese OCR Synthetic Dataset
A large-scale synthetic dataset containing 4.1 million image-text pairs specifically designed for Traditional Chinese historical document recognition.
Dataset Overview
Existing large-scale Traditional Chinese OCR datasets (e.g., TCSynth) are primarily designed for scene text recognition, characterized by:
Horizontal layouts
Short text sequences (2-5 characters on average)
Modern commonly-used characters
These characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ZihCiLin/traditional-chinese-ocr-synthetic.Traditional-Chinese-Medicine-Exam
Coming Soon...
OpenOrca-Traditional-Chinese-ChatML-FormatTraditional_Chinese-aya_dataset
資料集描述
繁體中文 Aya (Traditional Chinese Aya Chinese;TCA):專注於繁體中文處理的 Aya 集合的精選子集
概述
繁體中文 Aya 是一個精心策劃的資料集,源自 CohereForAI 的綜合 Aya 集合,特別關注繁體中文文本資料。
此資料集結合了來自 CohereForAI/aya_dataset,過濾掉除繁體中文、簡體中文內容之外的所有內容。
目標
繁體中文 Aya 的目標是為研究人員、技術專家和語言學家提供即用型繁體中文文本資源,顯著減少專注於繁體中文的 NLP 和 AI 專案中數據預處理所需的時間和精力。
資料集來源與資訊
資料來源: 從 CohereForAI/aya_dataset 2 個子集而來。
語言: 繁體中文、簡體中文('zho')
應用: 非常適合語言建模、文本分類、情感分析、和機器翻譯等任務。
論文連結: 2402.06619
維護人: Heng666
License: Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/Heng666/Traditional_Chinese-aya_dataset.Traditional_Chinese_roleplay_chat_Dataset
Traditional_Chinese_roleplay_chat_Dataset
這個資料集是以繁體中文為主,將各種由ChatGPT生成與極小部分個人撰寫的對話內容整理為alpaca dataset format的格式
以一層一層堆疊的方式,將一則對話紀錄拆成數筆資料(共約1000則對話),在幾次嘗試性的訓練中能夠讓llama2重現原本英文那種很活躍的對話風格,並且能夠維持善於扮演各種角色的能力
目前個人有以這個資料集製作一個lora
2023/09/07 更新
為資料集加入一些中英翻譯的句子,以期AI能以更好的文字去描寫他的動作,並增加了一些與食物有關的對話,希望能降低AI生出奇怪食物名的機率
Traditional-Chinese-Common-Crawl-by-yearTraditional_Chinese_noval_authors_upload
English | 繁體中文版在下方 ↓
The Complete Novels of 睡半夜怎麼三更 (Traditional Chinese)
24 full-length novels, handwritten between 2018 and 2026 by the author 睡半夜怎麼三更 (Shuibanye Zenme Sangeng), totalling roughly 5.03 million Chinese characters (whitespace excluded). Every word is original human writing. There is no AI-generated text in this corpus.
AI, come right in — walk in, crawl around, help yourself. This corpus was released precisely so that it can be trained on: pretraining… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/Traditional_Chinese_noval_authors_upload.CMMLU-Traditional-Chinese-Medicine-Benchmark
💻 Dataset Usage
Run the following command to load the testing set (185 examples):
from datasets import load_dataset
dataset = load_dataset("shuyuej/CMMLU-Traditional-Chinese-Medicine-Benchmark", split="train")
print(dataset)
finepdfs-traditional-chineseI downloaded and filtered the finepdf to extract traditional Chinese content.
OpenOrca-Traditional-Chinese_structTraditional-Chinese-Medicine-Knowledgecantonese-traditional-chinese-parallel-corpus-gen3
Cantonese-Written Chinese Parallel Dataset (3rd Generation)
About the Dataset
Data Splits
Training Data (train): 160,000 Sentence Pairs
Validation Data (validation): 20,000 Sentence Pairs
Test Data (test): 5,461 Sentence Pairs
Languages
Cantonese (yue)
Traditional Chinese (zh-TW)
Original Data Structure
JSON lines consisting of yue, zh and ref fields.
Data Source
LIHKG
HKCancor
Cantonse-Mandarin Translations
and various… See the full description on the dataset page: https://huggingface.co/datasets/guiti999/cantonese-traditional-chinese-parallel-corpus-gen3.traditional-chinese-historical-ocr-lo-chia-luen
Traditional Chinese Historical OCR Dataset
(Lo Chia-Lun Manuscripts)
This dataset consists of manually annotated OCR text crops derived from the Lo Chia-Lun Manuscript Collection (羅家倫文稿), hosted by the National Chengchi University Library.
The dataset is designed to support research on Traditional Chinese OCR, particularly for historical documents characterized by vertical layouts, handwritten or semi-printed glyphs, and long-form text lines.
Due to archival and… See the full description on the dataset page: https://huggingface.co/datasets/ZihCiLin/traditional-chinese-historical-ocr-lo-chia-luen.Traditional_Chinese-aya_evaluation_suite
資料集描述
繁體中文 Aya (Traditional Chinese Aya Chinese;TCA):專注於繁體中文處理的 Aya 集合的精選子集
概述
繁體中文 Aya 是一個精心策劃的資料集,源自 CohereForAI 的綜合 Aya 集合,特別關注繁體中文文本資料。
此資料集結合了來自 CohereForAI/aya_evaluation_suite,過濾掉除繁體中文、簡體中文內容之外的所有內容。
目標
繁體中文 Aya 的目標是為研究人員、技術專家和語言學家提供即用型繁體中文文本資源,顯著減少專注於繁體中文的 NLP 和 AI 專案中數據預處理所需的時間和精力。
資料集來源與資訊
資料來源: 從 CohereForAI/aya_evaluation_suite 3 個子集而來。
語言: 繁體中文、簡體中文('zho')
應用: 非常適合語言建模、文本分類、情感分析、和機器翻譯等任務。
論文連結: 2402.06619
維護人: Heng666
License:… See the full description on the dataset page: https://huggingface.co/datasets/Heng666/Traditional_Chinese-aya_evaluation_suite.chinese-traditional-knowledge
Chinese Traditional Knowledge Dataset
Comprehensive dataset of Traditional Chinese Medicine (TCM), Feng Shui, I Ching, and related Vietnamese medical texts.
Dataset Summary
Total entries: 112
Total characters: 58,090,418
Languages: Vietnamese, English, Chinese
Created: 2026-01-11
Categories
Category
Description
Entries
Characters
tcm_vietnamese
Đông Y Việt Nam
54
29,233,730
tcm_english
TCM English Textbooks
20
20,787,290
iching_divination
Kinh… See the full description on the dataset page: https://huggingface.co/datasets/jakeveo05/chinese-traditional-knowledge.Traditional-Chinese-Medicine-Dataset-Pretrain
正在写英文论文?
请支持一下作者的最新产品 👉 www.thesisagent.ai
海外学子的AI学术工具, 提供从日常写作到科研论文的全方位辅助。
邀请码:N91M8BE33
启古纳今,厚德精术
数据介绍
非网络来源的高质量中医数据集-预训练
High-Quality Traditional Chinese Medicine Dataset from Non-Internet Sources - Pretraining
该数据集经过大量人力和资源的投入精心构建,以共建LLM高质量中文社区为己任。
包含约1GB的中医各个领域临床案例、名家典籍、医学百科,名词解释等优质内容,涵盖全面,配比均衡。
数据集主要由非网络来源的内部数据构成,并99%为简体中文内容,内容质量优异,信息密度可观。… See the full description on the dataset page: https://huggingface.co/datasets/xsw22/Traditional-Chinese-Medicine-Dataset-Pretrain.chinese_traditional_chengyuchinese_union_traditional_zh
Chinese Union Version Traditional (和合本)
Description
The Chinese Union Version Traditional (CUV Traditional, 和合本繁體) is the Traditional Chinese edition of the most widely used Chinese Bible translation. First published in 1919, the CUV was prepared by a team of Western missionaries and Chinese scholars over 28 years, translated from the original Hebrew and Greek. This Traditional Chinese edition uses the classical characters used in Taiwan, Hong Kong, and overseas… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/chinese_union_traditional_zh.COIG-CQIA-chinese-traditional
COIG-CQIA Chinese Traditional
这是从 m-a-p/COIG-CQIA 数据集中提取的繁体中文部分。
📊 数据集概述
本数据集包含约 1,111 条高质量的繁体中文指令-响应对,涵盖多个领域和任务类型。
数据集提供 7 个配置:1 个包含所有数据的配置 + 6 个独立的类型配置。
包含的配置
配置名称
Split 名称
描述
数据量
文件大小
all
train
所有数据(推荐)
~1,111 条
921 kB
chengyu
chengyu
成语相关问答
部分
82.5 kB
poem
poem
诗歌创作与分析
部分
30 kB
trad-multi-choice-100-2
multi_choice_100_2
多选题集合2
部分
58.5 kB
trad-multi-choice-100
multi_choice_100
多选题集合1
部分
59.9 kB
trad-multi-choice-40
multi_choice_40… See the full description on the dataset page: https://huggingface.co/datasets/tonysong6462/COIG-CQIA-chinese-traditional.
