datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
taigi-literature-asts
Dataset Summary
The dataset contains 2,494 rows.
These paragraphs are extracted from authorized novels written by Ang Siok Tsiau洪淑昭 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 2,494 (each representing a paragraph)
Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-asts.taigi-literature-ttshs
Dataset Summary
The dataset contains 240 rows.
These paragraphs are extracted from authorized novel written by Tiunn Tshing Siong張青松 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 240 (each representing a paragraph)
Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ttshs.taigi-literature-abt
Dataset Summary
The dataset contains 389 rows.
These paragraphs are extracted from authorized novels written by Ang Bing-To洪明道 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 389 (each representing a paragraph)
Features:
title: Book… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-abt.taigi-literature-ngkh
Dataset Summary
The dataset contains 980 rows.
These paragraphs are extracted from authorized paper written by Ngoo Ka Hun吳嘉芬 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 980 (each representing a paragraph)
Features:
title: Paper… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ngkh.taigi-literature-lgs
Dataset Summary
The dataset contains 190 rows.
These paragraphs are extracted from authorized novels written by Lua Giok Si賴玉絲 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 190 (each representing a paragraph)
Features:
title: Book… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-lgs.taigi-literature-llb
Dataset Summary
The dataset contains 1,827 rows.
These paragraphs are extracted from authorized novels written by Lîm lang-bín林央敏 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 1,827 (each representing a paragraph)
Features:
title:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-llb.taigi-literature-ots
Dataset Summary
The dataset contains 5,260 rows.
These paragraphs are extracted from authorized novels written by Ou Tiong Siong胡長松 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 5,260 (each representing a paragraph)
Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ots.taigi-literature-kkh
Dataset Summary
The dataset contains 125 rows.
These paragraphs are extracted from authorized novels written by Ko Ka-hui高嘉徽 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 125 (each representing a paragraph)
Features:
title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-kkh.taigi-literature-ljk
Dataset Summary
The dataset contains 579 rows.
These paragraphs are extracted from authorized paper written by Lin Jui-Kun林瑞崐 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 579 (each representing a paragraph)
Features:
title: Paper… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-ljk.taigi-literature-olbt
Dataset Summary
The dataset contains 617 rows.
These paragraphs are extracted from authorized novels written by Ong Lo-Bit-To王羅蜜多 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 617 (each representing a paragraph)
Features:
title:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-olbt.taigi-literature-tks
Dataset Summary
The dataset contains 2,591 rows.
These paragraphs are extracted from authorized novels written by Tan Kim-Sun陳金順 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 2,591 (each representing a paragraph)
Features:
title: Book… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-tks.taigi-literature-tsk
Dataset Summary
The dataset contains 84 rows.
These paragraphs are extracted from authorized novels written by Tan Siu Ki陳秀枝 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 84 (each representing a paragraph)
Features:
title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-tsk.taigi-literature-sslts
Dataset Summary
The dataset contains 443 rows.
These paragraphs are extracted from authorized novels written by Sio Siann Ling Tsi小城綾子 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 443 (each representing a paragraph)
Features:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-sslts.taigi-literature-achiak
Dataset Summary
The dataset contains 609 rows.
These paragraphs are extracted from authorized prose written by Liau Tiunn Chiak廖張皭 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 609 (each representing a paragraph)
Features:
title:… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-achiak.taigi-literature-khg
Dataset Summary
The dataset contains 377 rows.
These paragraphs are extracted from authorized novels written by Khng Guan康原 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 377 (each representing a paragraph)
Features:
title: Book title… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-khg.taigi-literature-pikh
Dataset Summary
The dataset contains 3,585 rows.
These paragraphs are extracted from authorized literary works written by Png Iau Khian方耀乾 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Dataset Structure
Number of rows: 3,585 (each representing a paragraph)… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-pikh.taigi-literature-manlajo
Dataset Summary
The dataset contains 8,490 rows of Taiwanese Taigi translations.
This collection features Manlajo's translations into Taiwanese Taigi from a diverse range of sources, including:
Classic Western literature (such as "Robinson Crusoe" and "The Old Man and the Sea")
Japanese literature (including haiku poetry)
Chinese-language works and stories
Humorous anecdotes and jokes
Various other texts from different cultural traditions
All translations maintain the literary… See the full description on the dataset page: https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-manlajo.English_Taigi_Dict
