datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zh-tw-wikipedia
台灣正體中文維基百科 (zh-tw Wikipedia)
截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。
A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py.
於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。
For development… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/zh-tw-wikipedia.Wikipedia-EN-FA-Accessibility-Bridge
Wikipedia EN-FA Accessibility Bridge
Current, attributable English and Persian Wikipedia article snapshots for pages
created during a 69-day recency window, plus an EN↔FA
counterpart index and static accessibility signals.
The reproducible full baseline is the official 2026-08-01 Wikimedia dump:
6,289,549 English articles without Persian, 129,821 Persian
articles without English, and 22,277,907 namespace-0 pages in the
combined parity index. Redirects are retained in the parity… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge.wikipedia-zh-mnbvc
zhwiki-mnbvc
分项目:爬取并处理中文维基百科语料
数据时间:202302-202305 (持续更新)
主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC
该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1
并且使用组员开发的去重工具进行数据格式化。
总行数(样本): 10,754,146
一个示例:
{
"文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt",
"是否待查文件": false,
"是否重复文件": false,
"文件大小": 558,
"simhash": 14363740497821204542,
"最长段落长度": 142,
"段落数": 6,
"去重段落数": 6,
"低质量段落数": 0,
"段落": [
{… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.zh-tw-wikipedia
台灣正體中文維基百科 (zh-tw Wikipedia)
截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。
A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py.
於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。
For development… See the full description on the dataset page: https://huggingface.co/datasets/sirzmkk/zh-tw-wikipedia.wikipedia-zh-742M
Dataset Card for lianghsun/wikipedia-zh
以繁體中文(zh-tw)語系為主的 Wikipedia 資料集。
Dataset Details
Dataset Description
本資料集由自行開發的爬蟲抓取 Wikipedia 上標註繁體中文語系(zh-tw)的文本內容,以確保文本語系是繁體中文。目前 Hugging Face 上標註為繁體中文語系的許多 Wikipedia 資料集,其實並非真正的 Wikipedia 資料集,而是來自 Wikimedia 的內容。本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。
為便於訓練,同一個 Wikipedia 頁面的語料已被切分為多個子語料,使用者可依需求進行合併處理。好比同為「亞瑟·柯南·道爾」的文本:
...
{"text": "同樣在1887年的南海城,他受到了樸茨茅斯文學與哲學學會(Portsmouth Literary and… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-742M.Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.wikipedia-RuDataset
AI Platforms Wikipedia RU Dataset
Большой русскоязычный энциклопедический корпус на основе материалов Wikipedia.
Датасет предназначен для экспериментов с continued pretraining, языковой адаптацией, retrieval-корпусами и оценкой локальных русскоязычных LLM.
Состав
Репозиторий содержит:
wikipedia_rudataset.parquet — русскоязычный Wikipedia-корпус в Parquet.
Ориентировочный размер: миллионы текстовых записей (1M<n<10M).
Назначение
continued pretraining /… See the full description on the dataset page: https://huggingface.co/datasets/aiplatforms/wikipedia-RuDataset.Persian-Wikipedia-Corpus
Overview
This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library.
Usage
from datasets import load_dataset
dataset = load_dataset("codersan/Persian-Wikipedia-Corpus")
Persian-Wikipedia-Corpus
A complete copy of Persian Wikimedia pages, The dataset contains articles… See the full description on the dataset page: https://huggingface.co/datasets/codersan/Persian-Wikipedia-Corpus.wikipedia_2003
Wikipedia-2003
Original dump: https://dumps.wikimedia.org/archive/2003/2003-05-16
This is a filtered and cleaned version of the 2003 Wikipedia dump.
Stats
Language
Size
Lines
Bosnian (bs)
77.6KB
78
Czech (cs)
392.8KB
354
Danish (da)
4.9MB
11,561
German (de)
23.47MB
18,490
English (en)
249MB
128,198
Esperanto (eo)
7.9MB
7,202
Spanish (es)
7.33MB
4,651
French (fr)
13.2MB
10,957
Croatian (hr)
1.2KB
3
Dutch (nl)
10.9MB
7,116
Polish (pl)… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/wikipedia_2003.wikipedia_multiple_choice_qa
Galician and Portuguese Multiple-Choice QA Instruction Subsets
Dataset description
This dataset contains two instruction-tuning subsets for multiple-choice question answering in Galician and Portuguese:
gl_wikipedia_multiple_choice_qa (1,486 instances)
pt_wikipedia_multiple_choice_qa (547 instances)
Both subsets are reformatted versions of QA data originally included in the cpt_instruction_datasets collection, adapted here as standalone instruction-style datasets.
Each… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/wikipedia_multiple_choice_qa.wikipedia-geotagged
Geotagged Wikipedia
Every Wikipedia article that carries coordinates, with its text.
from datasets import load_dataset
ds = load_dataset("yuiseki/wikipedia-geotagged", "20260901.en")
ds = load_dataset("yuiseki/wikipedia-geotagged", "20260901.ja")
subset
articles
characters
share of the wiki
20260901.en
1,374,056
4,331,110,851
19.0% of 7,235,024
20260901.ja
218,496
435,046,691
14.4% of 1,516,331
Subsets are named {dump}.{lang}, as in
wikimedia/wikipedia.
A… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/wikipedia-geotagged.oman-wikipedia-corpus
Oman Wikipedia Corpus | مجموعة ويكيبيديا العُمانية
A bilingual (Arabic + English) plaintext corpus of Wikipedia articles about Oman,
built from the official MediaWiki APIs of ar.wikipedia.org and en.wikipedia.org.
مجموعة نصوص ثنائية اللغة (العربية والإنجليزية) من مقالات ويكيبيديا المتعلقة بسلطنة عُمان،
مبنية من واجهات ميدياويكي الرسمية.
Dataset Summary
Arabic (عربي)
English
Articles
1,205
522
Total words
702,327
400,257
Categories walked
258
147… See the full description on the dataset page: https://huggingface.co/datasets/jaifar/oman-wikipedia-corpus.wikipedia-20250721Dataset Card for Wikipedia-20250721
Dataset Summary
Wikipedia-20250721 is a cleaned, preprocessed version of the English Wikipedia “pages-articles” dump (July 21, 2025) converted into Parquet format and published on Hugging Face. It contains article titles and full article text suitable for large-scale language model pretraining or downstream NLP tasks.
Homepage: https://huggingface.co/datasets/yash3056/wikipedia-20250721
Dataset license: CC BY-SA 4.0
Languages: English
Size:… See the full description on the dataset page: https://huggingface.co/datasets/yash3056/wikipedia-20250721.ru-wikipedia-100k-full-text-daily-stats-10-years
📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews
**Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025**
📖 Описание
Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет.
Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.Persian-Wikipedia-Corpus
Overview
This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library.
Usage
from datasets import load_dataset
dataset = load_dataset("codersan/Persian-Wikipedia-Corpus")
Persian-Wikipedia-Corpus
A complete copy of Persian Wikimedia pages, The dataset contains articles… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/Persian-Wikipedia-Corpus.wikipedia-zh-filtered
Dataset Card for wikipedia-zh-filtered
本資料集是 lianghsun/wikipedia-zh-742M(中文維基百科語料)的繁體中文(zh-tw)過濾子集,每筆樣本含維基條目段落、token/字元統計與原始 URL。可作為繁中模型在百科類知識上的預訓練語料。
Dataset Details
Dataset Description
原始 wikipedia-zh-742M 涵蓋多種中文(簡體、繁體);本資料集刻意過濾出以繁體(zh-tw)為主的條目段落,並排除被識別為簡體為主的樣本。內容涵蓋臺灣本土條目(地理、歷史、人物)以及通用百科知識。
每筆樣本欄位:
text:條目內文段落。
token_count / word_count:token 數與字元數。
url:原條目網址(部分樣本可能為空字串)。
updated_at:抓取/更新時間戳。
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-filtered.arabic-wikipedia-wikibooks-corpus
Arabic Wiki Corpus - Parquet Format
Dataset Description
This dataset contains CLEANED Arabic text from Wikipedia and Wikibooks in Parquet format (efficient, fast, columnar storage).
Key Features
✅ Parquet format - faster loading, smaller size, columnar storage
✅ Fully cleaned - no markup, no HTML, no references
✅ Arabic normalized - alef, ya, ta marbuta normalized
✅ Ready for ML/NLP/LLM - use directly without preprocessing
Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-wikipedia-wikibooks-corpus.Persian-Wikipedia-Corpus
Overview
This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library.
Usage
from datasets import load_dataset
dataset = load_dataset("codersan/Persian-Wikipedia-Corpus")
Persian-Wikipedia-Corpus
A complete copy of Persian Wikimedia pages, The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AAkhoram/Persian-Wikipedia-Corpus.zh-tw-wikipedia
台灣正體中文維基百科 (zh-tw Wikipedia)
截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。
A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py.
於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。
For… See the full description on the dataset page: https://huggingface.co/datasets/mindbound0/zh-tw-wikipedia.spai-ss6-corpus-thai-wikipedia-clean
SPAI SS6 Thai Wikipedia Clean Corpus Index
Index repo for the Thai Wikipedia clean corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_wikipedia_clean_20230101
Rows in canonical config: 1,436,054
Parquet size in canonical config: 0.26 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-wikipedia-clean.wikipedia-pt-br-instructions
wikipedia-pt-br-instructions-gemma
Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR.
Origem
Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0.
Processo
A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wikipedia-pt-br-instructions.wikipedia-pt-br-instructions
wikipedia-pt-br-instructions-gemma
Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR.
Origem
Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0.
Processo
A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.
