CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zetavg /zh-tw-wikipedia 台灣正體中文維基百科 (zh-tw Wikipedia) 截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。 A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py. 於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。 For development… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/zh-tw-wikipedia.tabulartext-generation1M<n<10M30 likes871 downloads3y agoHugging Face02Reza2kn /Wikipedia-EN-FA-Accessibility-Bridge Wikipedia EN-FA Accessibility Bridge Current, attributable English and Persian Wikipedia article snapshots for pages created during a 69-day recency window, plus an EN↔FA counterpart index and static accessibility signals. The reproducible full baseline is the official 2026-08-01 Wikimedia dump: 6,289,549 English articles without Persian, 129,821 Persian articles without English, and 22,277,907 namespace-0 pages in the combined parity index. Redirects are retained in the parity… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge.tabulartranslation10M<n<100M1 likes661 downloads1mo agoHugging Face03wanng /wikipedia-zh-mnbvc zhwiki-mnbvc 分项目:爬取并处理中文维基百科语料 数据时间:202302-202305 (持续更新) 主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC 该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1 并且使用组员开发的去重工具进行数据格式化。 总行数(样本): 10,754,146 一个示例: { "文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt", "是否待查文件": false, "是否重复文件": false, "文件大小": 558, "simhash": 14363740497821204542, "最长段落长度": 142, "段落数": 6, "去重段落数": 6, "低质量段落数": 0, "段落": [ {… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.tabulartext-generation1M<n<10M6 likes390 downloads3y agoHugging Face04sirzmkk /zh-tw-wikipedia 台灣正體中文維基百科 (zh-tw Wikipedia) 截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。 A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py. 於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。 For development… See the full description on the dataset page: https://huggingface.co/datasets/sirzmkk/zh-tw-wikipedia.tabulartext-generation1M<n<10M0 likes343 downloads5mo agoHugging Face05lianghsun /wikipedia-zh-742M Dataset Card for lianghsun/wikipedia-zh 以繁體中文(zh-tw)語系為主的 Wikipedia 資料集。 Dataset Details Dataset Description 本資料集由自行開發的爬蟲抓取 Wikipedia 上標註繁體中文語系(zh-tw)的文本內容,以確保文本語系是繁體中文。目前 Hugging Face 上標註為繁體中文語系的許多 Wikipedia 資料集,其實並非真正的 Wikipedia 資料集,而是來自 Wikimedia 的內容。本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。 為便於訓練,同一個 Wikipedia 頁面的語料已被切分為多個子語料,使用者可依需求進行合併處理。好比同為「亞瑟·柯南·道爾」的文本: ... {"text": "同樣在1887年的南海城,他受到了樸茨茅斯文學與哲學學會(Portsmouth Literary and… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-742M.tabulartext-generation1M<n<10M4 likes222 downloads2y agoHugging Face06CJJones /Wikipedia_RAG_QA_Classification 🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training 📊 Dataset Description This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning. 🖥️ Demo Interface: Discord Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.tabulartext-generation100K<n<1M1 likes113 downloads6mo agoHugging Face07aiplatforms /wikipedia-RuDataset AI Platforms Wikipedia RU Dataset Большой русскоязычный энциклопедический корпус на основе материалов Wikipedia. Датасет предназначен для экспериментов с continued pretraining, языковой адаптацией, retrieval-корпусами и оценкой локальных русскоязычных LLM. Состав Репозиторий содержит: wikipedia_rudataset.parquet — русскоязычный Wikipedia-корпус в Parquet. Ориентировочный размер: миллионы текстовых записей (1M<n<10M). Назначение continued pretraining /… See the full description on the dataset page: https://huggingface.co/datasets/aiplatforms/wikipedia-RuDataset.tabulartext-generation1M<n<10M4 likes99 downloads5mo agoHugging Face08codersan /Persian-Wikipedia-Corpus Overview This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library. Usage from datasets import load_dataset dataset = load_dataset("codersan/Persian-Wikipedia-Corpus") Persian-Wikipedia-Corpus A complete copy of Persian Wikimedia pages, The dataset contains articles… See the full description on the dataset page: https://huggingface.co/datasets/codersan/Persian-Wikipedia-Corpus.tabulartext-generation1M<n<10M5 likes81 downloads2y agoHugging Face09fromziro /wikipedia_2003 Wikipedia-2003 Original dump: https://dumps.wikimedia.org/archive/2003/2003-05-16 This is a filtered and cleaned version of the 2003 Wikipedia dump. Stats Language Size Lines Bosnian (bs) 77.6KB 78 Czech (cs) 392.8KB 354 Danish (da) 4.9MB 11,561 German (de) 23.47MB 18,490 English (en) 249MB 128,198 Esperanto (eo) 7.9MB 7,202 Spanish (es) 7.33MB 4,651 French (fr) 13.2MB 10,957 Croatian (hr) 1.2KB 3 Dutch (nl) 10.9MB 7,116 Polish (pl)… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/wikipedia_2003.tabulartext-generation100K<n<1M0 likes80 downloads2mo agoHugging Face10proxectonos /wikipedia_multiple_choice_qa Galician and Portuguese Multiple-Choice QA Instruction Subsets Dataset description This dataset contains two instruction-tuning subsets for multiple-choice question answering in Galician and Portuguese: gl_wikipedia_multiple_choice_qa (1,486 instances) pt_wikipedia_multiple_choice_qa (547 instances) Both subsets are reformatted versions of QA data originally included in the cpt_instruction_datasets collection, adapted here as standalone instruction-style datasets. Each… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/wikipedia_multiple_choice_qa.tabulartext-generation1K<n<10K1 likes55 downloads5mo agoHugging Face11yuiseki /wikipedia-geotagged Geotagged Wikipedia Every Wikipedia article that carries coordinates, with its text. from datasets import load_dataset ds = load_dataset("yuiseki/wikipedia-geotagged", "20260901.en") ds = load_dataset("yuiseki/wikipedia-geotagged", "20260901.ja") subset articles characters share of the wiki 20260901.en 1,374,056 4,331,110,851 19.0% of 7,235,024 20260901.ja 218,496 435,046,691 14.4% of 1,516,331 Subsets are named {dump}.{lang}, as in wikimedia/wikipedia. A… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/wikipedia-geotagged.tabulartext-generation100K<n<1M1 likes53 downloads9h agoHugging Face12jaifar /oman-wikipedia-corpus Oman Wikipedia Corpus | مجموعة ويكيبيديا العُمانية A bilingual (Arabic + English) plaintext corpus of Wikipedia articles about Oman, built from the official MediaWiki APIs of ar.wikipedia.org and en.wikipedia.org. مجموعة نصوص ثنائية اللغة (العربية والإنجليزية) من مقالات ويكيبيديا المتعلقة بسلطنة عُمان، مبنية من واجهات ميدياويكي الرسمية. Dataset Summary Arabic (عربي) English Articles 1,205 522 Total words 702,327 400,257 Categories walked 258 147… See the full description on the dataset page: https://huggingface.co/datasets/jaifar/oman-wikipedia-corpus.tabulartext-generation1K<n<10K0 likes41 downloads3d agoHugging Face13yash3056 /wikipedia-20250721Dataset Card for Wikipedia-20250721 Dataset Summary Wikipedia-20250721 is a cleaned, preprocessed version of the English Wikipedia “pages-articles” dump (July 21, 2025) converted into Parquet format and published on Hugging Face. It contains article titles and full article text suitable for large-scale language model pretraining or downstream NLP tasks. Homepage: https://huggingface.co/datasets/yash3056/wikipedia-20250721 Dataset license: CC BY-SA 4.0 Languages: English Size:… See the full description on the dataset page: https://huggingface.co/datasets/yash3056/wikipedia-20250721.tabulartext-generation1M<n<10M1 likes34 downloads1y agoHugging Face14Mikimi /ru-wikipedia-100k-full-text-daily-stats-10-years 📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews **Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025** 📖 Описание Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет. Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.tabulartext-generation10K<n<100K1 likes27 downloads9mo agoHugging Face15mainkilora /Persian-Wikipedia-Corpus Overview This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library. Usage from datasets import load_dataset dataset = load_dataset("codersan/Persian-Wikipedia-Corpus") Persian-Wikipedia-Corpus A complete copy of Persian Wikimedia pages, The dataset contains articles… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/Persian-Wikipedia-Corpus.tabulartext-generation1M<n<10M1 likes16 downloads8mo agoHugging Face16lianghsun /wikipedia-zh-filteredgated Dataset Card for wikipedia-zh-filtered 本資料集是 lianghsun/wikipedia-zh-742M(中文維基百科語料)的繁體中文(zh-tw)過濾子集,每筆樣本含維基條目段落、token/字元統計與原始 URL。可作為繁中模型在百科類知識上的預訓練語料。 Dataset Details Dataset Description 原始 wikipedia-zh-742M 涵蓋多種中文(簡體、繁體);本資料集刻意過濾出以繁體(zh-tw)為主的條目段落,並排除被識別為簡體為主的樣本。內容涵蓋臺灣本土條目(地理、歷史、人物)以及通用百科知識。 每筆樣本欄位: text:條目內文段落。 token_count / word_count:token 數與字元數。 url:原條目網址(部分樣本可能為空字串)。 updated_at:抓取/更新時間戳。 Curated by: Huang Liang Hsun Language(s) (NLP): Traditional… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-filtered.tabulartext-generation10K<n<100K0 likes15 downloads5mo agoHugging Face17ArabicNLPWorld /arabic-wikipedia-wikibooks-corpusgated Arabic Wiki Corpus - Parquet Format Dataset Description This dataset contains CLEANED Arabic text from Wikipedia and Wikibooks in Parquet format (efficient, fast, columnar storage). Key Features ✅ Parquet format - faster loading, smaller size, columnar storage ✅ Fully cleaned - no markup, no HTML, no references ✅ Arabic normalized - alef, ya, ta marbuta normalized ✅ Ready for ML/NLP/LLM - use directly without preprocessing Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-wikipedia-wikibooks-corpus.tabulartext-generation1M<n<10M1 likes15 downloads4mo agoHugging Face18AAkhoram /Persian-Wikipedia-Corpus Overview This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library. Usage from datasets import load_dataset dataset = load_dataset("codersan/Persian-Wikipedia-Corpus") Persian-Wikipedia-Corpus A complete copy of Persian Wikimedia pages, The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AAkhoram/Persian-Wikipedia-Corpus.tabulartext-generation1M<n<10M0 likes13 downloads2mo agoHugging Face19mindbound0 /zh-tw-wikipedia 台灣正體中文維基百科 (zh-tw Wikipedia) 截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。 A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py. 於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。 For… See the full description on the dataset page: https://huggingface.co/datasets/mindbound0/zh-tw-wikipedia.tabulartext-generation1M<n<10M0 likes12 downloads2mo agoHugging Face20SPAISS6F1 /spai-ss6-corpus-thai-wikipedia-clean SPAI SS6 Thai Wikipedia Clean Corpus Index Index repo for the Thai Wikipedia clean corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: thai_wikipedia_clean_20230101 Rows in canonical config: 1,436,054 Parquet size in canonical config: 0.26 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-wikipedia-clean.tabulartext-generationn<1K0 likes9 downloads4mo agoHugging Face21br-llm-data /wikipedia-pt-br-instructionsgated wikipedia-pt-br-instructions-gemma Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR. Origem Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0. Processo A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wikipedia-pt-br-instructions.tabulartext-generationn<1K0 likes6 downloads3mo agoHugging Face22costadev00 /wikipedia-pt-br-instructionsgated wikipedia-pt-br-instructions-gemma Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR. Origem Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0. Processo A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.tabulartext-generationn<1K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.