CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lianghsun /finetranslations-edu-zhtwgated Dataset Card for finetranslations-edu-zhtw ✅ 狀態:翻譯已完成 —— 上游 HuggingFaceFW/finetranslations-edu 的 108,882,733 列已全數翻譯完成(逐檔清點兩邊 parquet footer,列數完全相符),共 8,195 個分片、約 1.26 TB。 ⚠️ 但尚未經過人工抽查驗證。翻譯全程自動化,使用前請自行評估品質。 📖 finetranslations-edu-zhtw 是以 HuggingFaceFW/finetranslations-edu 為來源,將其 translated_chunks(原始多語言教育類網頁內容、先被 pivot 翻譯成英文的版本)進一步翻譯成繁體中文的資料集。 Dataset Details Dataset Description HuggingFaceFW/finetranslations-edu 收錄了原本以英文以外語言(og_language,涵蓋約 200… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/finetranslations-edu-zhtw.tabulartext-generation100M<n<1B2 likes6.3k downloads2d agoHugging Face02zetavg /zh-tw-wikipedia 台灣正體中文維基百科 (zh-tw Wikipedia) 截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。 A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py. 於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。 For development… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/zh-tw-wikipedia.tabulartext-generation1M<n<10M30 likes911 downloads3y agoHugging Face03lianghsun /ultra-fineweb-l3-en-zhtwgated Ultra-FineWeb-L3 (en, multi_style) → 繁體中文 把 openbmb/Ultra-FineWeb-L3 的 英文 multi_style 子集(多風格改寫的預訓練語料)翻譯成台灣繁體中文。 欄位 欄位 說明 uid 來源自帶的唯一識別碼,未重算,可直接對回來源 style 來源的 style 欄位(此子集皆為 multi_style) content 英文原文 content_zhtw 繁體中文譯文 chunk_count 翻譯時切成幾塊(長文本會在段落邊界切塊後分別翻譯再接合) translation_error 該列是否有任何一塊翻譯失敗 翻譯規範 譯文遵循以下規則產生: 一律使用台灣繁體中文與台灣慣用詞,禁止簡體字與中國大陸慣用語 標點採全形/半形混排:中文語句用全形標點,英文單字、專有名詞、網址、數字維持半形 長文本在段落邊界切塊後分別翻譯再接合(chunk_count 記錄切了幾塊)… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/ultra-fineweb-l3-en-zhtw.texttext-generation10M<n<100M0 likes895 downloads5h agoHugging Face04erhwenkuo /c4-chinese-zhtw Dataset Card for "c4-chinese-zhtw" 內容 Common Crawl 是一個非營利組織,負責抓取網路並向公眾免費提供其檔案和資料集。Common Crawl 的網路檔案包含自 2008 年以來收集的 PB 級資料。它一般每月完成一次抓取。 Common Crawl 的爬蟲程式遵守 nofollow 和 robots.txt 政策。用於處理 Common Crawl 資料集的開源程式碼是公開可用的。 這個繁中的數據來是來自 Common Crawl 2023-14 的 data archive 下載并進行清理 。 這是 jed351 準備的版本,託管在這個位址: https://huggingface.co/datasets/jed351/Traditional-Chinese-Common-Crawl-Filtered 支援的任務 C4主要用於預訓練語言模型(pretrain language model)。 範例 一個樣本的範例: {… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/c4-chinese-zhtw.texttext-generation1M<n<10M12 likes714 downloads3y agoHugging Face05erhwenkuo /zhwikisource-zhtw Dataset Card for "zhwikisource-zhtw" 維基文庫(英文:Wikisource), 又稱 "自由的圖書館", 是一個由志願者在線收集自由內容文本的站點。 它屬維基媒體計劃項目,由維基媒體基金會負責運營。 作品類型: 典籍 | 史書 | 小說 | 詩歌 | 散文 | 演講 | 歌詞 | 經書 | 更多…… 主題: 條約 | 憲法 | 法律 | 教育 | 政治 | 歷史 | 宗教 | 更多…… 精選: 文章: 道德經 | 脂硯齋重評石頭記 文集: 紅樓夢 | 三國演義 | 西遊記 | 詩經 | 夢溪筆談 | 三十六計 | 古文觀止 歷史: 史記 | 資治通鑑 | 續資治通鑑 | 金史 | 漢書 | 後漢書 | 三國志 判例: 中國大理院解釋 | 中華民國最高法院解釋 | 中華民國司法院解釋 | 中華民國司法院大法官解釋 分類: 中華民國法律 | 中華人民共和國法律 | 中華人民共和國國務院政府工作報告 | 十三經 | 正史 這個數據集是根據 Wikipedia dumps… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/zhwikisource-zhtw.texttext-generation100K<n<1M4 likes658 downloads3y agoHugging Face06erhwenkuo /clean_passages_80m-chinese-zhtw Dataset Card for "clean_passages_80m-chinese-zhtw" 包含8千萬餘萬(88328203)個中文段落,不包含任何字母、數字。文字長度大部分介於 50~200 個字。 原始資料集是用於訓練GENIUS模型中文版。論文參考引用: @article{guo2022genius, title={GENIUS: Sketch-based Language Model Pre-training via Extreme and Selective Masking for Text Generation and Augmentation}, author={Guo, Biyang and Gong, Yeyun and Shen, Yelong and Han, Songqiao and Huang, Hailiang and Duan, Nan and Chen, Weizhu}, journal={arXiv preprint arXiv:2211.10330}, year={2022} }… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/clean_passages_80m-chinese-zhtw.texttext-generation10M<n<100M2 likes454 downloads3y agoHugging Face07sirzmkk /zh-tw-wikipedia 台灣正體中文維基百科 (zh-tw Wikipedia) 截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。 A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py. 於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。 For development… See the full description on the dataset page: https://huggingface.co/datasets/sirzmkk/zh-tw-wikipedia.tabulartext-generation1M<n<10M0 likes341 downloads5mo agoHugging Face08liswei /news-collection-zhtw Dataset Card for Traditional Chinese News Collection Contains common news/magazines/articles available online in Traditional Chinese. Provides title, text (content), and category for each sample. Note: category is labeled according to the source of the news. Cleaned with custom rules and de-deplicated using MinHash. Dataset Details Dataset size: 557,764 samples. Available labels: article tech science daily-weekly Dataset source: benchang1110/technewstw… See the full description on the dataset page: https://huggingface.co/datasets/liswei/news-collection-zhtw.texttext-generation100K<n<1M3 likes241 downloads2y agoHugging Face09liswei /common-crawl-zhtw Dataset Card for Common Crawl Traditional Chinese De-duplicated version of jed351/Traditional-Chinese-Common-Crawl-Filtered. De-duplicated with MinHash Is suggested to filter the dataset with NLU models before any serious use. texttext-generation1M<n<10M6 likes229 downloads2y agoHugging Face10erhwenkuo /openorca-chinese-zhtw Dataset Card for "openorca-chinese-zhtw" Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope. The data is primarily used for training and evaluation in the field of natural… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/openorca-chinese-zhtw.texttext-classification1M<n<10M3 likes228 downloads3y agoHugging Face11Blaze7451 /Wiki-zhtw-20250601 Dataset Card for Wiki-zhtw-20250601 Dataset Description This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC. texttext-generation1M<n<10M1 likes182 downloads1y agoHugging Face12liswei /c4-zhtw Dataset Card for C4-zhtw Traditional Chinese subset of the C4 dataset. De-duplicated with MinHash. Is suggested to filter the dataset with NLU models before any serious use. texttext-generation1M<n<10M3 likes179 downloads2y agoHugging Face13erhwenkuo /firefly-train-chinese-zhtw Dataset Card for "firefly-train-chinese-zhtw" 資料集摘要 本資料集主要是應用於專案:Firefly(流螢): 中文對話式大語言模型 ,經過訓練後得到的模型 firefly-1b4。 [Firefly(流螢): 中文對話式大語言模型]專案(https://github.com/yangjianxin1/Firefly)收集了 23 個常見的中文資料集,并且對於每種不同的 NLP 任務,由人工書寫若干種指令模板來保證資料的高品質與豐富度。 資料量為115萬 。數據分佈如下圖所示: 訓練資料集的 token 長度分佈如下圖所示,絕大部分資料的長度都小於 600: 原始資料來源: YeungNLP/firefly-train-1.1M Firefly(流萤): 中文对话式大语言模型 資料下載清理 下載 chinese-poetry: 最全中文诗歌古典文集数据库 的 Repo 使用 OpenCC 來進行簡繁轉換 使用 Huggingface Datasets 來上傳至… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/firefly-train-chinese-zhtw.texttext-generation1M<n<10M2 likes148 downloads3y agoHugging Face14asd567557275 /zhtw-roleplay-space-grimoire Space Grimoire RP Corpus (Traditional Chinese) Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0. 中文說明在下方 Dataset Summary Source text 283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.tabulartext-generation10K<n<100K1 likes132 downloads14d agoHugging Face15erhwenkuo /wikipedia-zhtw Dataset Card for "wikipedia-zhtw" 維基百科數據集包含許多不同語言的文章。這個數據集是根據 Wikipedia dumps (https://dumps.wikimedia.org/) 裡頭 zhwiki 的中文下載檔案來建構的。每個範例都包含一篇完整的維基百科文章的內容,並經過清理以去除不需要的部分(例如參考文獻等)。 Homepage: https://dumps.wikimedia.org zhwiki 下載點: https://dumps.wikimedia.org/zhwiki 數據 Dump 版本 由於維基百科數據集定期會進行網站數據拋轉,在 2023/10/10 的時間點去查看時會有下列的數據可供下載: 數據 Dump 目錄 拋轉時間點 20230620/ 01-Aug-2023 09:31 20230701/ 20-Aug-2023 09:41 20230720/ 01-Sep-2023 09:31 20230801/ 20-Sep-2023… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/wikipedia-zhtw.texttext-generation1M<n<10M7 likes122 downloads3y agoHugging Face16erhwenkuo /poetry-chinese-zhtw Dataset Card for "poetry-chinese-zhtw" 資料集摘要 中文古典文集資料庫收集了約 5.5 萬首唐詩、26 萬首宋詩、2.1 萬首宋詞和其他古典文集。詩人包括唐宋兩朝近 1.4 萬古詩人,和兩宋時期 1.5 千古詞人。 五代十國- 收錄"花間集"與"南唐二主詞" 唐- 收錄"全唐詩"(是清康熙四十四年,康熙皇帝主導下,蒐集羅唐詩的收藏「得詩 48,900 餘首,詩入 2,200 人」)。 宋- 收錄"全宋詞"(由唐圭璋編著,孔凡禮補輯,共收錄宋代詞人 1,330 家,詞作 21,116 首)。 元- 收錄元曲 11,057 篇,曲家 233 人。 清- 收錄"納蘭性德詩集" 原始資料來源: chinese-poetry: 最全中文诗歌古典文集数据库 資料下載清理 下載 chinese-poetry: 最全中文诗歌古典文集数据库 的 Repo 調整資料呈現結構便於模型訓練 使用 OpenCC 來進行簡繁轉換 使用 Huggingface Datasets 來上傳至… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/poetry-chinese-zhtw.texttext-generation10K<n<100K21 likes98 downloads3y agoHugging Face17erhwenkuo /medical_dialogue-chinese-zhtw Dataset Card for "medical_dialogue-chinese-zhtw" 中文醫療問答資料集 來源 本資料集是從 Toyhom/Chinese-medical-dialogue-data 的 github repo 中轉換而來。 內容 科別 數量 Andriatria 男科 94,596 個問答對 IM 內科 220,606 個問答對 OAGD 婦產科 183,751 個問答對 Oncology 腫瘤科 75,553 個問答對 Pediatric 兒科 101,602 個問答對 Surgical 外科 115,991 個問答對 總計 792,099 條數據 範例 { "instruction": "現在你是個神經腦外科醫生,請根據病人的問題給予建議:", "input": "癲癇病能吃德巴金嗎,錯覺,有時候感覺看到的和聽到的不太一樣。", "output":… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/medical_dialogue-chinese-zhtw.texttext-generation100K<n<1M25 likes74 downloads3y agoHugging Face18erhwenkuo /train_2m-chinese-zhtw Dataset Card for "train_2m-chinese-zhtw" 內容 包含約 200 萬條由 BELLE 專案目產生的中文指令(instruction)資料。 範例 { "instruction": "將以下三個句子組合成一個有意義的段落。\n狗是人類最好的朋友。它們非常聰明,可以進行各種活動。如果你喜歡散步,狗可以成為你一起散步的夥伴。", "input": "", "output": "狗是人類最好的朋友,它們非常聰明,可以進行各種活動。如果你喜歡散步,狗可以成為你一起散步的伙伴。出門散步是一種良好的鍛煉方式,而有狗的陪伴會讓散步變得更有趣,並且有狗在身邊也能給你帶來安全感。所以,擁有一隻狗作為你的伙伴,可以幫助你變得更加積極主動和健康。" } 欄位: instruction: 指令 input: 輸入(此資料集均為空) output: 輸出 使用限制… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/train_2m-chinese-zhtw.texttext-generation1M<n<10M2 likes65 downloads3y agoHugging Face19liswei /wikipedia-zhtw-dedupDeduplicate version of erhwenkuo/wikipedia-zhtw using MinHash. texttext-generation1M<n<10M4 likes58 downloads2y agoHugging Face20DataAgent /TCNNet-SFT-NetCom-zhTW-1.1Mgated [TCNNet] A Traditional Chinese Networking and Communication Instruction Fine-Tuning Dataset (zh-TW) A large-scale supervised fine-tuning (SFT) dataset created specifically for TCNNet-9B, a Chinese language model specialized in networking and communications domains. The dataset contains question-answer pairs generated from various networking, cybersecurity, and tech review articles written in Traditional Chinese. Dataset Description Dataset Summary This dataset… See the full description on the dataset page: https://huggingface.co/datasets/DataAgent/TCNNet-SFT-NetCom-zhTW-1.1M.texttext-generation1K<n<10K3 likes57 downloads2y agoHugging Face21lianghsun /fineweb-zhtwgated🍷 FineWeb-zhtw 是從 🍷 FineWeb 抽取繁體中文內容所建立的全量、未經品質篩選網頁語料集,共 48,058,113 列。 它是 📚 FineWeb-Edu-zhtw 的上游來源——後者以教育導向分類器從本資料集篩選而成。 Dataset Details Dataset Description 🍷 FineWeb-zhtw 取自 🍷 FineWeb(第一代)中判定為繁體中文的網頁文本,保留原始的 text、id 與抓取階段的中繼資料,未做教育性或品質篩選。 原始 FineWeb 涵蓋龐大的 Common Crawl 網頁,其中繁體中文的部分同樣夾雜廣告、論壇、導覽列與低資訊密度內容。本資料集的定位是提供完整、未經取捨的基底,讓後續的篩選方法本身可被研究與比較;若需要可直接訓練的高品質語料,請改用 📚 FineWeb-Edu-zhtw。 列數 48,058,113 檔案 322 個 parquet 大小 約 99.5 GB Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/fineweb-zhtw.texttext-generation10M<n<100M0 likes55 downloads1mo agoHugging Face22erhwenkuo /multiturn_chat_0.8m-chinese-zhtw Dataset Card for "multiturn_chat_0.8m-chinese-zhtw" 內容 包含約 80 萬條由 BELLE 專案所產生的 Human 與 Assistant 的多輪對話。 注意:此資料集是由 ChatGPT 產生的,未經嚴格校驗,內容可能包含錯誤。使用過程中請注意這一點。 instruction 中包含多輪對話的上文內容,以 Human: 和 Assistant: 區分,output 中包含當前 Assistant 角色的回答。 範例 { "instruction": "Human:請根據題材「健身房」,寫一篇800字的文章。 \nAssistant:好的,以下是由我生成的800字「健身房」主題文章:\n在現代社會,人們越來越珍惜健康和身材。健身房就是一個人們可以鍛鍊身體的場所,也是一個社交的場所。在健身房,人們可以透過器材訓練,有氧運動和瑜珈等方式來達到健身的目的。… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/multiturn_chat_0.8m-chinese-zhtw.texttext-generation100K<n<1M7 likes49 downloads3y agoHugging Face23liswei /coct-en-zhtw-dedup Dataset Card for COCT en-zhtw deduplicate De-duplicate version of zetavg/coct-en-zh-tw-translations-twp-300k. * The du-duplicate process is run twice for both English and Traditional Chinese. texttranslation100K<n<1M3 likes47 downloads2y agoHugging Face24erhwenkuo /alpaca-data-gpt4-chinese-zhtw Dataset Card for "alpaca-data-gpt4-chinese-zhtw" This dataset contains Chinese (zh-tw) Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This dataset is a translation from English to Chinese. Dataset structure It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca. The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/alpaca-data-gpt4-chinese-zhtw.texttext-generation10K<n<100K6 likes45 downloads3y agoHugging Face25AWeirdDev /zh-tw-pts-articles-sm zh-tw-pts-articles-sm 🐣English • 🇹🇼 繁體中文 This dataset contains articles scraped from PNN News. It's a news provider verified by the vast majority. Note: some keys like conclusion may be None. Dataset({ features: ['image', 'title', 'conclusion', 'content', 'timestamp', 'category', 'link'], num_rows: 1400 }) Use The Dataset Use 🤗 Datasets to download, use or modify this dataset. from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/AWeirdDev/zh-tw-pts-articles-sm.imagetext-generation1K<n<10K7 likes44 downloads3y agoHugging Face26renhehuang /coffee-order-zhtw 本資料集由OllaForge生成 Coffee Order Dataset (繁體中文/Traditional Chinese) 專為咖啡點餐場景設計的繁體中文多輪對話資料集,適用於訓練任務導向對話系統。 資料集描述 此資料集包含模擬咖啡店點餐場景的多輪對話,涵蓋各種點餐情境,包括: 基本點餐流程 訂單修改與取消 處理菜單外品項 處理超出限制的請求(如加兩份濃縮) 口語化表達理解 語言 繁體中文(台灣) 包含台灣口語表達(如「ㄋㄟㄋㄟ」) 資料規模 項目 數量 總對話數 2,939 平均輪數 2-6 輪 語言 繁體中文 資料格式 每筆資料為 JSONL 格式,包含 conversations 欄位: { "conversations": [ { "role": "system", "content":… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/coffee-order-zhtw.texttext-generation1K<n<10K0 likes41 downloads9mo agoHugging Face27AWeirdDev /zh-tw-essays zh-tw-essays (12K) Essays obtained from 勵志人生 - Zeelive. from datasets import load_dataset dataset = load_dataset("AWeirdDev/zh-tw-essays") Format { "title": "孩子童年不吃苦,家長晚年必吃苦" # The title "link": "https://www.zeelive.com.tw/jiatingjiaoyu/184191.html", "content": "錢財莫輕,勤苦得來;奢華莫學,自取貧窮…" # Text content. **May be blank!** } texttext-generation10K<n<100K1 likes40 downloads2y agoHugging Face28AWeirdDev /zh-tw-recipes-smtexttext-generation1K<n<10K2 likes38 downloads3y agoHugging Face29erhwenkuo /wikinews-zhtw Dataset Card for "wikinews-zhtw" 維基新聞(英文:Wikinews)是由一群志願者、即民間記者運營的網路媒體。同時是一個自由內容的維基,屬維基媒體計劃項目,由維基媒體基金會負責運營。維基新聞通過協作新聞學的工作模式去運行,同時亦努力通過中性的觀點報導新聞,包括原創一手獨家報道和採訪。 這個數據集是根據 Wikipedia dumps (https://dumps.wikimedia.org/) 裡頭 zhwikinews 的中文下載檔案來建構的。每個範例都包含一篇完整的維基新聞文章的內容,並經過清理以去除不需要的部分。 Homepage: https://dumps.wikimedia.org zhwiki 下載點: https://dumps.wikimedia.org/zhwikinews 數據 Dump 版本 由於維基百科數據集定期會進行網站數據拋轉,在 2023/10/10 的時間點去查看時會有下列的數據可供下載: 數據 Dump 目錄 拋轉時間點… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/wikinews-zhtw.texttext-generation1K<n<10K6 likes37 downloads3y agoHugging Face30liswei /wikinews-zhtw-dedupDeduplicate version of erhwenkuo/wikinews-zhtw using MinHash. texttext-generation1K<n<10K0 likes34 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.