CoolFace
Datasetpublic

lianghsun/wikipedia-zh-742M

Dataset Card for lianghsun/wikipedia-zh 以繁體中文(zh-tw)語系為主的 Wikipedia 資料集。 Dataset Details Dataset Description 本資料集由自行開發的爬蟲抓取 Wikipedia 上標註繁體中文語系(zh-tw)的文本內容,以確保文本語系是繁體中文。目前 Hugging Face 上標註為繁體中文語系的許多 Wikipedia 資料集,其實並非真正的 Wikipedia 資料集,而是來自 Wikimedia 的內容。本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。 為便於訓練,同一個 Wikipedia 頁面的語料已被切分為多個子語料,使用者可依需求進行合併處理。好比同為「亞瑟·柯南·道爾」的文本: ... {"text": "同樣在1887年的南海城,他受到了樸茨茅斯文學與哲學學會(Portsmouth Literary and… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-742M.

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
4likes223downloads
Dataset Card

Dataset Card for lianghsun/wikipedia-zh

<!-- Provide a quick summary of the dataset. --> 以繁體中文(zh-tw)語系為主的 Wikipedia 資料集。

Dataset Details

Dataset Description

<!-- Provide a longer summary of what this dataset is. --> 本資料集由自行開發的爬蟲抓取 Wikipedia 上標註繁體中文語系(zh-tw)的文本內容,以確保文本語系是繁體中文。目前 Hugging Face 上標註為繁體中文語系的許多 Wikipedia 資料集,其實並非真正的 Wikipedia 資料集,而是來自 Wikimedia 的內容。本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。

為便於訓練,同一個 Wikipedia 頁面的語料已被切分為多個子語料,使用者可依需求進行合併處理。好比同為「亞瑟·柯南·道爾」的文本:

yaml
...
{"text": "同樣在1887年的南海城,他受到了樸茨茅斯文學與哲學學會(Portsmouth Literary and Philosophical Society)會員阿爾弗雷德·威爾克斯·德雷森(英語:Alfred Wilks Drayson)的影響,他開始一系列的超自然調查。其中包括參加約20次的降神會、心靈感應實驗和與靈媒混在一起。他寫給唯靈論雜誌《Light》,宣稱自己是一名唯靈論者,並且談到一件令他信服的特殊超自然事件。", "key": "亞瑟·柯南·道爾", "url": "https://zh.wikipedia.org/zhtw/%E9%98%BF%E7%91%9F%C2%B7%E6%9F%AF%E5%8D%97%C2%B7%E9%81%93%E7%88%BE", "word_count": 210, "token_count": "", "updated_date": "2024-12-05"}
{"text": "雖然後來他動搖過,但他仍然對超自然現象著迷。1889年,他成為漢普郡心靈研究學會(Hampshire Society for Psychical Research)的創始人之一,並於1893年加入了倫敦的心靈研究學會(英語:Society for Psychical Research)。他在1894年加入了Sidney Scott爵士和法蘭克·波德莫爾(英語:Frank Podmore),在德文郡進行了一次鬧鬼調查。儘管如此,在此期間,他本質上仍是一名業餘。", "key": "亞瑟·柯南·道爾", "url": "https://zh.wikipedia.org/zh-tw/%E9%98%BF%E7%91%9F%C2%B7%E6%9F%AF%E5%8D%97%C2%B7%E9%81%93%E7%88%BE", "word_count": 231, "token_count": "", "updated_date": "2024-12-05"}
{"text": "其後柯南·道爾的妻子路易斯·霍金斯、兒子Kingsley Doyle、弟弟Innes、兩個姐夫及侄子相繼去世,令道爾寄情於唯靈論之中從而獲得慰藉,他亦試圖想證明可以與死者的靈魂溝通,例如通靈。他也曾經加入了一個研究超自然的協會「鬼魂俱樂部」。", "key": "亞瑟·柯南·道爾", "url": "https://zh.wikipedia.org/zh-tw/%E9%98%BF%E7%91%9F%C2%B7%E6%9F%AF%E5%8D%97%C2%B7%E9%81%93%E7%88%BE", "word_count": 121, "token_count": "", "updated_date": "2024-12-05"}
...

本資料集經 [meta-llama/Llama-3.2-1B]() 統計共有 742,565,363 tokens。

附註:爬蟲程式中斷在某個時間,故資料集可能尚未包含全部繁體中文的內容,將不定期更新此資料集。

  • —Curated by: Huang Liang Hsun
  • —Language(s) (NLP): Tranditional Chinese
  • —License: cc-by-nc-sa-4.0

Dataset Sources [optional]

<!-- Provide the basic links for the dataset. -->

Uses

<!-- Address questions around how the dataset is intended to be used. -->

Direct Use

<!-- This section describes suitable use cases for the dataset. -->

[More Information Needed]

Out-of-Scope Use

<!-- This section addresses misuse, malicious use, and uses that the dataset will not work well for. --> 再次強調:本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。

Dataset Structure

<!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. -->

[More Information Needed]

Dataset Creation

Curation Rationale

<!-- Motivation for the creation of this dataset. -->

[More Information Needed]

Source Data

<!-- This section describes the source data (e.g. news text and headlines, social media posts, translated sentences, ...). -->

Data Collection and Processing

<!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. -->

[More Information Needed]

Who are the source data producers?

<!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. --> Huang Liang Hsun

Annotations [optional]

<!-- If the dataset contains annotations which are not part of the initial data collection, use this section to describe them. -->

Annotation process

<!-- This section describes the annotation process such as annotation tools used in the process, the amount of data annotated, annotation guidelines provided to the annotators, interannotator statistics, annotation validation, etc. -->

[More Information Needed]

Who are the annotators?

<!-- This section describes the people or systems who created the annotations. -->

[More Information Needed]

Personal and Sensitive Information

<!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. -->

[More Information Needed]

Bias, Risks, and Limitations

<!-- This section is meant to convey both technical and sociotechnical limitations. -->

再次強調:本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。

Recommendations

<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->

Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.

Citation [optional]

<!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. -->

BibTeX:

[More Information Needed]

APA:

[More Information Needed]

Glossary [optional]

<!-- If relevant, include terms and calculations in this section that can help readers understand the dataset or dataset card. -->

[More Information Needed]

More Information [optional]

[More Information Needed]

Dataset Card Authors

Huang Liang Hsun

Dataset Card Contact

Huang Liang Hsun