CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lemon07r /bartowski-imatrix-v5-semantic Bartowski iMatrix Calibration v5 (Semantic Chunking) A processed version of bartowski's v5 imatrix calibration data using semantic boundary detection optimized for the v5 data structure. Dataset Summary Metric Value Total samples 2,075 Chunking method V5-optimized semantic boundary detection Chunk size 200+ characters (no upper limit, preserves document integrity) Languages English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic.texttext-generation1K<n<10K9 likes158 downloads8mo agoHugging Face02lemon07r /bartowski-imatrix-v3-semantic Bartowski iMatrix Calibration v3 (Semantic Chunking) A processed version of bartowski's v3 imatrix calibration data using semantic boundary detection in attempt to create coherent, non-overlapping samples. Dataset Summary Metric Value Total samples 168 Chunking method Semantic boundary detection Target chunk size ~2048 characters Languages English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese Source Data The… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v3-semantic.texttext-generationn<1K1 likes33 downloads8mo agoHugging Face03IMA-Taiwan /taigi-literature-astsgated Dataset Summary The dataset contains 2,494 rows. These paragraphs are extracted from authorized novels written by Ang Siok Tsiau洪淑昭 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 2,494 (each representing a paragraph) Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-asts.text1K<n<10K0 likes13 downloads1y agoHugging Face04IMA-Taiwan /ima-corpus-zhtwgated IMA Traditional Chinese Corpus(繁體中文語料總集) 本資料集為繁體中文文學語料總集,目的在於將原先分散於多個作者/來源 dataset repo 的繁體中文文本統一整併,提供「一次申請、持續更新」的集中存取方式。 使用者只需申請本 dataset(本 repo)一次,即可取得所有繁中語料。未來新增來源或更新資料將直接同步至本 repo,無需重複申請。 📂 目錄結構 所有來源資料皆保留於 data/ 之下,每個子資料夾對應一個原始來源 repo,例如: data/ ├── zhtw-literature-ots 每個子資料夾內保留: 原始 README 原始語料檔(json / txt 等) 來源資訊與授權說明 以利來源追溯與資料審核。 📦 資料格式 主要格式: JSON UTF-8 編碼文字檔 典型欄位可能包含: title:作品名稱 author:作者 content:文本內容 source:來源 repo (依各來源資料實際格式而定)… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/ima-corpus-zhtw.texttext-generation1K<n<10K0 likes12 downloads8mo agoHugging Face05IMA-Taiwan /taigi-literature-ttshsgated Dataset Summary The dataset contains 240 rows. These paragraphs are extracted from authorized novel written by Tiunn Tshing Siong張青松 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 240 (each representing a paragraph) Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ttshs.textn<1K0 likes11 downloads1y agoHugging Face06IMA-Taiwan /zhtw-literature-otsgated Dataset Summary The dataset contains 2,349 rows. These paragraphs are extracted from authorized novels written by Ou Tiong Siong胡長松 and contain multiple sentences in Traditional Chinese. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 2,349 (each representing a paragraph) Features: title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/zhtw-literature-ots.text1K<n<10K1 likes11 downloads1y agoHugging Face07IMA-Taiwan /taigi-literature-abtgated Dataset Summary The dataset contains 389 rows. These paragraphs are extracted from authorized novels written by Ang Bing-To洪明道 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 389 (each representing a paragraph) Features: title: Book… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-abt.textn<1K0 likes10 downloads2y agoHugging Face08IMA-Taiwan /taigi-literature-ngkhgated Dataset Summary The dataset contains 980 rows. These paragraphs are extracted from authorized paper written by Ngoo Ka Hun吳嘉芬 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 980 (each representing a paragraph) Features: title: Paper… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ngkh.textn<1K0 likes10 downloads1y agoHugging Face09IMA-Taiwan /taigi-literature-llbgated Dataset Summary The dataset contains 1,827 rows. These paragraphs are extracted from authorized novels written by Lîm lang-bín林央敏 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 1,827 (each representing a paragraph) Features: title:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-llb.text1K<n<10K0 likes10 downloads9mo agoHugging Face10IMA-Taiwan /taigi-literature-otsgated Dataset Summary The dataset contains 5,260 rows. These paragraphs are extracted from authorized novels written by Ou Tiong Siong胡長松 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 5,260 (each representing a paragraph) Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ots.text1K<n<10K6 likes9 downloads1y agoHugging Face11IMA-Taiwan /taigi-literature-kkhgated Dataset Summary The dataset contains 125 rows. These paragraphs are extracted from authorized novels written by Ko Ka-hui高嘉徽 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 125 (each representing a paragraph) Features: title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-kkh.textn<1K0 likes9 downloads1y agoHugging Face12IMA-Taiwan /taigi-literature-ljkgated Dataset Summary The dataset contains 579 rows. These paragraphs are extracted from authorized paper written by Lin Jui-Kun林瑞崐 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 579 (each representing a paragraph) Features: title: Paper… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ljk.textn<1K0 likes9 downloads2y agoHugging Face13IMA-Taiwan /taigi-literature-olbtgated Dataset Summary The dataset contains 617 rows. These paragraphs are extracted from authorized novels written by Ong Lo-Bit-To王羅蜜多 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 617 (each representing a paragraph) Features: title:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-olbt.textn<1K0 likes9 downloads2y agoHugging Face14IMA-Taiwan /taigi-literature-lgsgated Dataset Summary The dataset contains 190 rows. These paragraphs are extracted from authorized novels written by Lua Giok Si賴玉絲 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 190 (each representing a paragraph) Features: title: Book… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-lgs.textn<1K0 likes9 downloads1y agoHugging Face15IMA-Taiwan /taigi-literature-tksgated Dataset Summary The dataset contains 2,591 rows. These paragraphs are extracted from authorized novels written by Tan Kim-Sun陳金順 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 2,591 (each representing a paragraph) Features: title: Book… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-tks.text1K<n<10K0 likes8 downloads2y agoHugging Face16IMA-Taiwan /taigi-literature-tskgated Dataset Summary The dataset contains 84 rows. These paragraphs are extracted from authorized novels written by Tan Siu Ki陳秀枝 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 84 (each representing a paragraph) Features: title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-tsk.textn<1K0 likes8 downloads1y agoHugging Face17IMA-Taiwan /taigi-literature-khggated Dataset Summary The dataset contains 377 rows. These paragraphs are extracted from authorized novels written by Khng Guan康原 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 377 (each representing a paragraph) Features: title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-khg.textn<1K0 likes8 downloads1y agoHugging Face18IMA-Taiwan /taigi-literature-ssltsgated Dataset Summary The dataset contains 443 rows. These paragraphs are extracted from authorized novels written by Sio Siann Ling Tsi小城綾子 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 443 (each representing a paragraph) Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-sslts.textn<1K0 likes7 downloads1y agoHugging Face19IMA-Taiwan /taigi-literature-achiakgated Dataset Summary The dataset contains 609 rows. These paragraphs are extracted from authorized prose written by Liau Tiunn Chiak廖張皭 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 609 (each representing a paragraph) Features: title:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-achiak.textn<1K0 likes7 downloads1y agoHugging Face20IMA-Taiwan /taigi-literature-pikhgated Dataset Summary The dataset contains 3,585 rows. These paragraphs are extracted from authorized literary works written by Png Iau Khian方耀乾 and contain multiple sentences in Taiwanese Taigi written with Hanji. The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research. Dataset Structure Number of rows: 3,585 (each representing a paragraph)… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-pikh.text1K<n<10K0 likes7 downloads1y agoHugging Face21IMA-Taiwan /taigi-literature-manlajogated Dataset Summary The dataset contains 8,490 rows of Taiwanese Taigi translations. This collection features Manlajo's translations into Taiwanese Taigi from a diverse range of sources, including: Classic Western literature (such as "Robinson Crusoe" and "The Old Man and the Sea") Japanese literature (including haiku poetry) Chinese-language works and stories Humorous anecdotes and jokes Various other texts from different cultural traditions All translations maintain the literary… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-manlajo.text1K<n<10K0 likes6 downloads1y agoHugging Face22IMA-Taiwan /zhtw-literature-meowversegated Dataset Summary The dataset contains 14,256 rows. These paragraphs are extracted from authorized works written by Chang Luo-Miao張珞喵 and published on Vocus 方格子. Most articles are composed in Traditional Chinese, and many of them also carry parallel English, Japanese, German and Spanish versions of the same content within the same article, so the dataset is multilingual by design rather than translated afterwards. Roughly 43% of the paragraphs are Traditional Chinese, 54% English… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/zhtw-literature-meowverse.text10K<n<100K0 likes6 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.