lianghsun/fineweb-zhtw
🍷 FineWeb-zhtw 是從 🍷 FineWeb 抽取繁體中文內容所建立的全量、未經品質篩選網頁語料集,共 48,058,113 列。 它是 📚 FineWeb-Edu-zhtw 的上游來源——後者以教育導向分類器從本資料集篩選而成。 Dataset Details Dataset Description 🍷 FineWeb-zhtw 取自 🍷 FineWeb(第一代)中判定為繁體中文的網頁文本,保留原始的 text、id 與抓取階段的中繼資料,未做教育性或品質篩選。 原始 FineWeb 涵蓋龐大的 Common Crawl 網頁,其中繁體中文的部分同樣夾雜廣告、論壇、導覽列與低資訊密度內容。本資料集的定位是提供完整、未經取捨的基底,讓後續的篩選方法本身可被研究與比較;若需要可直接訓練的高品質語料,請改用 📚 FineWeb-Edu-zhtw。 列數 48,058,113 檔案 322 個 parquet 大小 約 99.5 GB Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/fineweb-zhtw.
021
No card is published for this repository, or it could not be fetched from Hugging Face right now.
