CoolFace
Datasetpublicgated

lianghsun/tw-news-551M

Dataset Card for tw-news-551M tw-news-551M 是一個台灣繁體中文之長文本預訓練語料集,包含 648,576 篇文章,總計約 5.63 億 tokens。涵蓋政治、社會、文化、環境、科技等多元主題,適用於語言模型在繁體中文語境之持續預訓練。 Dataset Details Dataset Description 本資料集彙整台灣繁體中文之公開文章,涵蓋政治、社會、文化、環境、科技等多元主題。每筆資料包含文章全文與 token/字數等統計元資料。 Curated by: Liang Hsun Huang Language(s) (NLP): Traditional Chinese License: CC BY-NC-SA 4.0 Dataset Sources Repository: lianghsun/tw-news-551M Uses Direct… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-news-551M.

sourceHugging Facecc-by-nc-sa-4.0updated 1mo agoView on Hugging Face
2likes14downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
lianghsun/tw-news-551M · CoolFace