datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-ja-20231030
Wikipedia Japanese data (20231030)
Source Date: 2023/10/30
Source: https://dumps.wikimedia.org/other/cirrussearch/
License
CC BY-SA 4.0
Example
WIP
wikipedia-paragraphs
wikipedia-paragraphs
wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research.
Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID.
Dataset structure
Configurations
The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.osm-polygon-wikidata-and-wikipedia
OSM Polygon Wikidata + Wikipedia, V2
V2 builds on the V1 Wikidata-only dataset with multilingual Wikipedia and Wikivoyage text. It also retains valid multilingual wikipedia=* references, including polygons without a Wikidata QID.
Source code: GitHub repository.
Dataset snapshot
Metric
Value
Polygon rows across regional extracts
1,259,424
Unique polygon identities (osm_type, osm_id)
1,188,854
Polygons with successful non-empty text (unique OSM… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-wikidata-and-wikipedia.wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2
Dataset Card for "wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2"
More Information needed
wikipedia_en
wikipedia_en
This is a curated Wikipedia English dataset for use with the II-Commons project.
Dataset Details
Dataset Description
This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.wikipedia-utils
Wikipedia-Utils: Preprocessed Wikipedia Texts for NLP
Preprocessed Wikipedia texts generated with the scripts in singletongue/wikipedia-utils repo.
For detailed information on how the texts are processed, please refer to the repo.
wikipedia-pageviews
Wikipedia Article Pageviews
This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia.
It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago.
The fetcher runs in a scheduled GitHub Actions workflow, which is available here.
The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.wikipedia-22-12-concat-split
Dataset Card for "wikipedia-22-12-concat-split"
More Information needed
zh-tw-wikipedia
台灣正體中文維基百科 (zh-tw Wikipedia)
截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。
A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py.
於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。
For development… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/zh-tw-wikipedia.wikipedia_fr_2022French part of Wikipedia embeded with Solon-embeddings-large-0.1.
wikipedia_2017
Dataset Card for Dataset Name
This is a Wikipedia dataset correct to "31-12-2017".
Dataset Details
Dataset Description
WikiMedia routinely publishes dumps of Wikipedia, each containing the revision history of articles. We first defined the relevant revision before extracting the article information. Specifically, we select the most recent revision as of December 31st for each year. Consequently, some revisions in our datasets date back several years from the… See the full description on the dataset page: https://huggingface.co/datasets/Ti-Ma/wikipedia_2017.Wikipedia-EN-FA-Accessibility-Bridge
Wikipedia EN-FA Accessibility Bridge
Current, attributable English and Persian Wikipedia article snapshots for pages
created during a 69-day recency window, plus an EN↔FA
counterpart index and static accessibility signals.
The reproducible full baseline is the official 2026-08-01 Wikimedia dump:
6,289,549 English articles without Persian, 129,821 Persian
articles without English, and 22,277,907 namespace-0 pages in the
combined parity index. Redirects are retained in the parity… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge.wikipedia_2018
Dataset Card for Dataset Name
This is a Wikipedia dataset correct to "31-12-2018".
Dataset Details
Dataset Description
WikiMedia routinely publishes dumps of Wikipedia, each containing the revision history of articles. We first defined the relevant revision before extracting the article information. Specifically, we select the most recent revision as of December 31st for each year. Consequently, some revisions in our datasets date back several years from the… See the full description on the dataset page: https://huggingface.co/datasets/Ti-Ma/wikipedia_2018.kNN-Targets-wikipedia-mistral
Dataset Overview
This dataset provides k-nearest neighbor (kNN) target distributions for language modeling. Each token in the Wikipedia corpus is associated with a soft probability distribution over its top-k nearest neighbors in the representation space of a frozen language model. These targets can be used to train MLP Memory.
Corresponding Preprocessed Corpus: Rubin-Wei/enwiki-dec2021-preprocessed-mistral
Compatible Model: Mistral-7B-v0.3
Paper: MLP Memory: A Retriever-Pretrained… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/kNN-Targets-wikipedia-mistral.Wikipedia-TR-2023-Embedded-Dump
Wikipedia-TR-2023-Embedded-Dump
Türkçe Vikipedi (tr.wikipedia.org) makaleleri, retrieval / RAG kullanım
senaryoları için parçalanmış (chunk) ve embedding'lenmiş hâliyle. Her
makale bir parent (ana) chunk'a (tüm makale metni, bağlam
genişletmek için) ve birden fazla child (alt) chunk'a (her biri
kendi embedding'ine sahip küçük pasajlar) bölünmüştür.
İçerik
Makale
348.751
Embedding'li child chunk
1.308.623
Parent chunk (embeddingsiz)
348.751… See the full description on the dataset page: https://huggingface.co/datasets/SalihHub/Wikipedia-TR-2023-Embedded-Dump.fineweb-2-spa-5m-wikipedia-2mfineweb-2-spa-3m-wikipedia-2mwikipedia-multilingual-synthetic-ir-query
wikipedia-multilingual-synthetic-ir-query
This dataset contains multilingual Wikipedia-derived synthetic query-document pairs for information retrieval training.
It was created with the query-crafter-multilingual model, which generates search-like queries from Wikipedia text.
The current release contains two different retrieval settings:
short_doc: pairs of (query, short document)
long_doc: pairs of (query, long document)
These two subsets are not generated in the same way… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-synthetic-ir-query.Multilingual-USAS-Labelled-Silver-Wikipedia
Multilingual USAS Silver Labelled Wikipedia Articles
Silver-labelled Wikipedia article text for training USAS semantic taggers and Multi-Word
Expression (MWE) identifiers, covering 8 Wikipedia language sites. The source text comes from the
HuggingFace HuggingFaceFW/finewiki
dataset, restricted to articles rated Good (GA) or Featured (FA) — using the article ID list from
ucrelnlp/wikipedia-ga-fa-ids — and
then sentence split and automatically tagged with USAS semantic tags and… See the full description on the dataset page: https://huggingface.co/datasets/ucrelnlp/Multilingual-USAS-Labelled-Silver-Wikipedia.wikipedia-zh-mnbvc
zhwiki-mnbvc
分项目:爬取并处理中文维基百科语料
数据时间:202302-202305 (持续更新)
主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC
该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1
并且使用组员开发的去重工具进行数据格式化。
总行数(样本): 10,754,146
一个示例:
{
"文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt",
"是否待查文件": false,
"是否重复文件": false,
"文件大小": 558,
"simhash": 14363740497821204542,
"最长段落长度": 142,
"段落数": 6,
"去重段落数": 6,
"低质量段落数": 0,
"段落": [
{… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.wikipedia-for-mask-filling\wikipedia-256-tokenThis is Wikidedia passages dataset for ODQA retriever.
Each passages have 256~ tokens splitteed by gpt-4 tokenizer using tiktoken.
Token count
{'~128': 1415068, '128~256': 1290011,
'256~512': 18756476, '512~1024': 667,
'1024~2048': 12, '2048~4096': 0, '4096~8192': 0,
'8192~16384': 0, '16384~32768': 0, '32768~65536': 0,
'65536~128000': 0, '128000~': 0}
Text count
{'~512': 1556876,'512~1024': 6074975, '1024~2048': 13830329,
'2048~4096': 49, '4096~8192': 2, '8192~16384': 3, '16384~32768': 0… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/wikipedia-256-token.zh-tw-wikipedia
台灣正體中文維基百科 (zh-tw Wikipedia)
截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。
A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py.
於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。
For development… See the full description on the dataset page: https://huggingface.co/datasets/sirzmkk/zh-tw-wikipedia.tokenized_wikipedia_20220301.en_train_512
Tokenized English Wikipedia Dataset
Dataset Description
This dataset contains tokenized chunks of text from the English Wikipedia dump of March 1, 2022. Each entry in the dataset represents a chunk of text from Wikipedia, with information about which document and position within the document it comes from.
Dataset Creation
Source Dataset: Wikipedia (20220301.en)
Tokenizer: BERT base uncased
Chunk Size: 512 tokens (including special tokens)… See the full description on the dataset page: https://huggingface.co/datasets/TemryL/tokenized_wikipedia_20220301.en_train_512.wikipedia_token
Dataset Card for "wikipedia_token"
Token count {
'~1024': 5320881,
'1024~2048': 693911,
'2048~4096': 300935,
'4096~8192': 106221,
'8192~16384': 30611,
'16384~32768': 4812,
'32768~65536': 1253,
'65536~128000': 46,
'128000~': 0
}
Text count {
'0~1024': 2751539,
'1024~2048': 1310778,
'2048~4096': 1179150,
'4096~8192': 722101,
'8192~16384': 329062,
'16384~32768': 121237,
'32768~65536': 36894,
'65536~': 7909
}
Token percent {
'~1024':… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/wikipedia_token.filtered_japanese-wikipediawikipedia_quality_wikirankDatasets with WikiRank quality score as of 1 August 2024 for 47 million Wikipedia articles in 55 language versions (also simplified version for each language in separate files).
The WikiRank quality score is a metric designed to assess the overall quality of a Wikipedia article. Although its specific algorithm can vary depending on the implementation, the score typically combines several key features of the Wikipedia article.
Why It’s Important
Enhances Trust: For readers and… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/wikipedia_quality_wikirank.Wikipedia-FA-EN-DeepSeek-V4-Flash-0731
Wikipedia Persian to English — DeepSeek V4 Flash 0731
Rolling, machine-generated English translations of Persian Wikipedia articles
from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration
full_articles_fa_without_en. 129,816 translations are
currently published in 26 immutable Parquet shards.
The target release contains 129,816 translations;
five source rows have empty plain_text and are not translated. Shards are
published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.wikipedia-qwen-4b-clustered-nodiskpq-7to8wikipedia-zh-742M
Dataset Card for lianghsun/wikipedia-zh
以繁體中文(zh-tw)語系為主的 Wikipedia 資料集。
Dataset Details
Dataset Description
本資料集由自行開發的爬蟲抓取 Wikipedia 上標註繁體中文語系(zh-tw)的文本內容,以確保文本語系是繁體中文。目前 Hugging Face 上標註為繁體中文語系的許多 Wikipedia 資料集,其實並非真正的 Wikipedia 資料集,而是來自 Wikimedia 的內容。本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。
為便於訓練,同一個 Wikipedia 頁面的語料已被切分為多個子語料,使用者可依需求進行合併處理。好比同為「亞瑟·柯南·道爾」的文本:
...
{"text": "同樣在1887年的南海城,他受到了樸茨茅斯文學與哲學學會(Portsmouth Literary and… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-742M.
