datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikimedia_filtered
Wikimedia
Description
Official Wikimedia wikis are released under a CC BY-SA license. We downloaded the official database dumps from March 2025 of the English-language wikis that are directly managed by the Wikimedia Foundation. These database dumps include the wikitext—MediaWiki’s custom markup language—for each page as well as talk pages, where editors discuss changes made to a page. We only use the most recent version of each page. We converted wikitext to plain text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikimedia_filtered.wikipedia-cn-20230720-filtered本数据集基于中文维基2023年7月20日的dump存档。作为一项以数据为中心的工作,本数据集仅保留了 254,547条 质量较高的词条内容。具体而言:
过滤了Template, Category, Wikipedia, File, Topic, Portal, MediaWiki, Draft, Help等特殊类型的词条
使用启发式的方法和自有的NLU模型过滤了一部分质量较低的词条
过滤了一部分内容较为敏感或存在争议性的词条。
进行了简繁转换和习惯用词转换,确保符合中国大陆地区的习惯用词。
This dataset is based on the Chinese Wikipedia dump archive from July 20th, 2023. As a data-centric effort, the dataset retains 254,574 high-quality entries. Specifically:
Entries of special types such as Template, Category, Wikipedia, File, Topic… See the full description on the dataset page: https://huggingface.co/datasets/0xDing/wikipedia-cn-20230720-filtered.wikiteam
Wikiteam
Description
There are many wikis on the internet that are not managed by the Wikimedia foundation, but do use their MediaWiki software to power their wiki.
Many of these wikis have been archived by wikiteam, a collection of volunteers that create unofficial database dumps of wikis and upload them to the Internet Archive.
We download all dumps made by wikiteam when the metadata indicates the wiki was licensed under CC BY, CC BY-SA, or released into the public… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikiteam.wikimedia
Wikimedia
Description
Official Wikimedia wikis are released under a CC BY-SA license.
We downloaded the official database dumps from March 2025 of the English-language wikis that are directly managed by the Wikimedia foundation.
These database dumps include the wikitext—Mediawiki’s custom markup language—for each page as well as talk pages, where editors discuss changes made for a page.
We only use the most recent version of each page.
We converted wikitext to plain… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikimedia.llm-jp-corpus-v4-ja_wiki
llm-jp-corpus-v4 — ja_wiki
Mirror of the ja/ja_wiki sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_wiki
Files: 6 × jsonl.gz (1.9 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY-SA 3.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki.wikipedia-paragraphs
Wikipedia Paragraph Samples
Dataset Description
This dataset contains paragraphs extracted from randomly selected English Wikipedia articles. It provides a diverse sample of Wikipedia content across various topics.
Dataset Details
Name: Wikipedia Paragraph Samples
Version: 1.0
Date Created: 2024-08-20
Language: English
Format: JSONLines
Contents
Each line in the dataset represents a single paragraph and contains two fields:
Title of the Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs.m2d2-wiki-decon
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/claran/m2d2-wiki-decon.wikipedia-zh-mnbvc
zhwiki-mnbvc
分项目:爬取并处理中文维基百科语料
数据时间:202302-202305 (持续更新)
主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC
该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1
并且使用组员开发的去重工具进行数据格式化。
总行数(样本): 10,754,146
一个示例:
{
"文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt",
"是否待查文件": false,
"是否重复文件": false,
"文件大小": 558,
"simhash": 14363740497821204542,
"最长段落长度": 142,
"段落数": 6,
"去重段落数": 6,
"低质量段落数": 0,
"段落": [
{… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.wikiteam_filtered
Wikiteam
Description
There are many wikis on the internet that are not managed by the Wikimedia Foundation, but do use their MediaWiki software to power their wiki. Many of these wikis have been archived by Wikiteam, a collection of volunteers that create unofficial database dumps of wikis and upload them to the Internet Archive. We download all dumps made by Wikiteam when the metadata indicates the wiki was licensed under CC BY, CC BY-SA, or released into the public… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikiteam_filtered.wikipedia-markdown
Marin Markdownified Wikipedia
Markdownified Wikipedia is a large-scale, pre-processed version of the English Wikipedia Enterprise HTML dump consisting of 8.59B tokens. The corpus has been converted to clean, section-aware Markdown for language-model training.
Value
Tokens
8 587 224 558
Primary source
https://dumps.wikimedia.org/other/enterprise_html/runs/20241201/enwiki-NS0-20241201-ENTERPRISE-HTML.json.tar.gz
File format
JSONL
License
CC-BY-SA 4.0 (mirrors… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/wikipedia-markdown.Wiki.korpus
Viki korpusi na srpskom i hrvatskom jeziku
Sveža verzija, 1. maj 2026!
Očišćen i filtriran skup pet projekata: Vikipedija, Vikizvornik, Vikiknjige, Vikivesti i Vikicitati.
Preko 670.000 očišćenih članaka, sa preko 310 miliona reči.
Svaki dokument je u zasebnoj JSON liniji.
Novi metapodaci! Kategorije, broj reči i postotak ćiriličnog teksta
Moguće filtiranje skupa po jeziku ili projektu.
Wiki corpora in Serbian and… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.korpus.danbooru-wiki-2026
danbooru-wiki-2026-04-28
About
Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag.
This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.WikiBigEdit
📚 WikiBigEdit
Paper (EasyEdit2 Framework): EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models Code (EasyEdit Framework): https://github.com/zjunlp/EasyEdit Project Page (EasyEdit2 Framework): https://zjunlp.github.io/project/EasyEdit2/ Paper (WikiBigEdit Dataset): Understanding the Limits of Lifelong Knowledge Editing in LLMs
🌟 Overview
WikiBigEdit is a large-scale benchmark designed to assess the ability of language models to integrate… See the full description on the dataset page: https://huggingface.co/datasets/lukasthede/WikiBigEdit.wikipedia-zh-742M
Dataset Card for lianghsun/wikipedia-zh
以繁體中文(zh-tw)語系為主的 Wikipedia 資料集。
Dataset Details
Dataset Description
本資料集由自行開發的爬蟲抓取 Wikipedia 上標註繁體中文語系(zh-tw)的文本內容,以確保文本語系是繁體中文。目前 Hugging Face 上標註為繁體中文語系的許多 Wikipedia 資料集,其實並非真正的 Wikipedia 資料集,而是來自 Wikimedia 的內容。本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。
為便於訓練,同一個 Wikipedia 頁面的語料已被切分為多個子語料,使用者可依需求進行合併處理。好比同為「亞瑟·柯南·道爾」的文本:
...
{"text": "同樣在1887年的南海城,他受到了樸茨茅斯文學與哲學學會(Portsmouth Literary and… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-742M.Wiki-zhtw-20250601
Dataset Card for Wiki-zhtw-20250601
Dataset Description
This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC.
verified_wiki_historian_the_beatles_anthology_dataset_active
Verified-Wiki-Historian: The Beatles Anthology
Verified-Wiki-Historian (The Beatles Anthology) is a refined, citation-grounded instruction dataset for Beatles-specific historical question answering, summarization, and supervised fine-tuning.
This dataset is a cleaned and rebuilt refinement of:
Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active
The current release contains 4,000 instruction records focused on Beatles history, recording sessions, release… See the full description on the dataset page: https://huggingface.co/datasets/Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active.wikireading
Dataset Card for Wikireading
This is a dataset of book chapters scraped from a Russian website called Wikireading.
Dataset Details
Dataset Description
Wikireading is a collection of non-fiction educational books in various domains: Biology, Art, History, Religion and much more. The books are highly educational and provide vast knowledge in different domains, making this dataset a good choice for pretraining.
The resulting dataset contains ~26M rows, which in… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/wikireading.wikipedia-first-paragraphwiki-to-rcqa-italian
Wiki-to-RCQA - Italian (IT)
openai-terra-batch-wiki-brazil-1000-partial-20260724-01
OpenAI Terra Batch — Wikipédia PT-BR (run parcial)
Checkpoint publicável de uma execução real e interrompida do fluxo
document_task_matrix. A execução planejou gerar uma matriz de 1.000
documentos da Wikipédia em português por 25 tasks canônicas usando a Responses
API Batch e o modelo gpt-5.6-terra.
Este repositório não representa a conclusão dos 25.000 pares planejados. Ele
contém somente os 1.282 candidatos aceitos após a reconciliação offline de
todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.wikipedia-human-ai
Wikipedia Human/AI
A selection of ~10,000 paragraphs from Wikipedia, along with rewritten text by GPT 5 Nano.
wiki_paragraphs_norwegian
WIKI Paragraphs Norwegian
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.Wiki.sl
Wiki korpusi v slovenščini
Sveža različica, 1. maj 2026!
Očiščen in filtriran nabor štirih projektov: Wikipedia, Wikivir, Wikiknjige in Wikinavedki.
Več kot 146.000 kuriranih člankov z več kot 130 milijoni besed.
Vsak dokument je v ločeni vrstici JSON.
Novi metapodatki! Kategorije, število besed (in odstotek ciriličnega besedila)
Možnost filtriranja nabora po jeziku ali projektu.
Wiki corpora in Slovenian language… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.sl.Wiki-ja-20250601
Dataset Card for Wiki-ja-20250601
Dataset Description
This dataset is derived from the Japan‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing.
imabari_wiki_qa_v4_validated_w_reasoning_effort_qwen38
Imabari Wiki QA v4 Validated with Reasoning Effort — Qwen3.8
概要 / Overview
日本語・今治弁のQAを用いて、reasoning effort に応じた思考文の生成を学習するための教師ありファインチューニング(SFT)用データセットです。Imabari Wiki QA v4 Validated の質問と回答を保持し、元記事の文脈を参照して思考文を再生成しています。
This dataset supports supervised fine-tuning (SFT) of reasoning-effort-conditioned explanations using Japanese QA with Imabari dialect expressions. Questions and answers from Imabari Wiki QA v4 Validated are preserved, while reasoning text is… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_validated_w_reasoning_effort_qwen38.imabari_wiki_qa_v4_validated
Imabari QA v4 — Validated
Dataset Summary
Imabari QA v4 — Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
This dataset combines two independently curated variants of the Imabari QA v4 synthetic reasoning dataset:
ikedachin/imabari_qa_v4_program_validated
ikedachin/imabari_qa_v4_human_validated
Both datasets are derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The combined… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_validated.Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.Wiki.mk
Вики корпуси на македонски јазик
Свежа верзија, 1 мај 2026!
Исчистен и филтриран сет од три проекти: Википедија, Викиизвор и Викикниги
Над 100.000 курирани статии, со над 50 милиони зборови.
Секој документ е во посебна JSON линија.
Нови метаподатоци! Категории, број на зборови и процент на кириличен текст
Можно е да се филтрира сет по јазик или проект.
Wiki corpora in Macedonian
Fresh version, 1. May 2026!… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.mk.Wiki-ko-20250601
Dataset Card for Wiki-ko-20250601
Dataset Description
This dataset is derived from the Korea‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing.
Wiki-vi-20250601
Dataset Card for Wiki-vi-20250601
Dataset Description
This dataset is derived from the Vietnam‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing.
