datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phantom-wiki-v0-5-0-predictions
Dataset Card for Dataset Name
Predictions from https://huggingface.co/datasets/mlcore/phantom-wiki-v050
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/phantom-wiki-v0-5-0-predictions.wiki-talks
Wiki-Talks
The Wiki-Talks dataset is a collection of conversational threads extracted from the talk pages on Wikipedia.
This dataset captures collaborative dialogue, discussion patterns, and consensus-building among Wikipedia contributors.
It is useful for NLP research focused on dialogue, sentiment analysis, and community dynamics.
Details
Currently due to PyArrow incompatibility to the long recursive structures in the dataset there is an intrinsic incompatibility… See the full description on the dataset page: https://huggingface.co/datasets/lflage/wiki-talks.wikipedia-zh-mnbvc
zhwiki-mnbvc
分项目:爬取并处理中文维基百科语料
数据时间:202302-202305 (持续更新)
主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC
该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1
并且使用组员开发的去重工具进行数据格式化。
总行数(样本): 10,754,146
一个示例:
{
"文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt",
"是否待查文件": false,
"是否重复文件": false,
"文件大小": 558,
"simhash": 14363740497821204542,
"最长段落长度": 142,
"段落数": 6,
"去重段落数": 6,
"低质量段落数": 0,
"段落": [
{… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.Wiki.korpus
Viki korpusi na srpskom i hrvatskom jeziku
Sveža verzija, 1. maj 2026!
Očišćen i filtriran skup pet projekata: Vikipedija, Vikizvornik, Vikiknjige, Vikivesti i Vikicitati.
Preko 670.000 očišćenih članaka, sa preko 310 miliona reči.
Svaki dokument je u zasebnoj JSON liniji.
Novi metapodaci! Kategorije, broj reči i postotak ćiriličnog teksta
Moguće filtiranje skupa po jeziku ili projektu.
Wiki corpora in Serbian and… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.korpus.danbooru-wiki-2026
danbooru-wiki-2026-04-28
About
Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag.
This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.wikipedia-zh-742M
Dataset Card for lianghsun/wikipedia-zh
以繁體中文(zh-tw)語系為主的 Wikipedia 資料集。
Dataset Details
Dataset Description
本資料集由自行開發的爬蟲抓取 Wikipedia 上標註繁體中文語系(zh-tw)的文本內容,以確保文本語系是繁體中文。目前 Hugging Face 上標註為繁體中文語系的許多 Wikipedia 資料集,其實並非真正的 Wikipedia 資料集,而是來自 Wikimedia 的內容。本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。
為便於訓練,同一個 Wikipedia 頁面的語料已被切分為多個子語料,使用者可依需求進行合併處理。好比同為「亞瑟·柯南·道爾」的文本:
...
{"text": "同樣在1887年的南海城,他受到了樸茨茅斯文學與哲學學會(Portsmouth Literary and… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-742M.wikipedia-qwen-4b-clustered-nodiskpq-7to8wiki-to-rcqa-italian
Wiki-to-RCQA - Italian (IT)
wikipedia_en_512_for_pretraining
Cleaned Wikipedia 512 Pretraining Dataset
Dataset Description
This dataset is a cleaned version of the Hugging Face dataset lucadiliello/wikipedia_512_pretraining.
The original dataset contains English Wikipedia text prepared for language-model pretraining.
This derivative version applies additional filtering intended to remove malformed, duplicated, or incomplete training samples while otherwise preserving the original text.
Source
Original… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/wikipedia_en_512_for_pretraining.wikiMIA-2024-hard
WikiMIA-2024 Hard Dataset
Dataset Description
WikiMIA_2024 Hard is a challenging dataset for membership inference attacks intorduced in the paper "The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage" containing temporal Wikipedia articles with different versions based on date cutoffs.
This dataset is designed to evaluate the robustness of privacy-preserving machine learning models against sophisticated membership inference techniques.
It… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/wikiMIA-2024-hard.Wiki.sl
Wiki korpusi v slovenščini
Sveža različica, 1. maj 2026!
Očiščen in filtriran nabor štirih projektov: Wikipedia, Wikivir, Wikiknjige in Wikinavedki.
Več kot 146.000 kuriranih člankov z več kot 130 milijoni besed.
Vsak dokument je v ločeni vrstici JSON.
Novi metapodatki! Kategorije, število besed (in odstotek ciriličnega besedila)
Možnost filtriranja nabora po jeziku ali projektu.
Wiki corpora in Slovenian language… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.sl.wiki_paragraphs_norwegian
WIKI Paragraphs Norwegian
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.german-wikipedia-clean-2phantom-wiki-v0.3-null
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Format based on this dataset: https://huggingface.co/rag-datasets
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/phantom-wiki-v0.3-null.Wiki.mk
Вики корпуси на македонски јазик
Свежа верзија, 1 мај 2026!
Исчистен и филтриран сет од три проекти: Википедија, Викиизвор и Викикниги
Над 100.000 курирани статии, со над 50 милиони зборови.
Секој документ е во посебна JSON линија.
Нови метаподатоци! Категории, број на зборови и процент на кириличен текст
Можно е да се филтрира сет по јазик или проект.
Wiki corpora in Macedonian
Fresh version, 1. May 2026!… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.mk.Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.wiki_paragraphs_english
WIKI Paragraphs English
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.qfs-smollm2-135m-wikitext2-native-v1
HF workflow d3dc69602aeb981f06bd9f4c726937f9
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.nor_wiki_reasoning_english_v4Wiki.bg
Уики корпуси на македонски и български език
Нова версия, 1 май 2026 г!
Почистен и филтриран набор от пет проекта: Уикипедия, Уикиизточник, Уикикниги, Уикиновини и Уикицитат.
Над 240 000 курирани статии, с над 100 милиона думи.
Всеки документ е на отделен JSON ред.
Нови метаданни! Категории, брой думи и процент на кирилица
Възможно е да филтрирате набора по език или проект.
Wiki corpora in Bulgarian
Fresh… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.bg.glm51-kld-reference-logits-wikitext-ctx2048-s512-20260517 ---
license: other
pretty_name: GLM-5.1 KLD Reference Logits WikiText ctx2048 s512
tags:
- logits
- kld
- glm-5.1
- vllm
- b12x
---
# GLM-5.1 KLD Reference Logits
Public cache of the reference logits used for GLM-5.1 NVFP4 / mixed
FP8_PB_WO KLD evaluation. These files are generated logits, not model
weights. They are stored as `logits_*.safetensors` with one tensor named
`logits`, shape `(2047, 154880)`, dtype `float32`.
##… See the full description on the dataset page: https://huggingface.co/datasets/festr2/glm51-kld-reference-logits-wikitext-ctx2048-s512-20260517.wikiSQL-kk-datasetsoyjak-wiki
Soyjak Wiki
A full dump of Soyjak Wiki, a MediaWiki-based encyclopedia documenting soyjak memes, variants, culture, communities, and related internet history. The dump includes all 8,039 pages (2,927 articles and 5,112 redirects) with raw wikitext markup preserved.
Columns
Column
Type
Description
title
string
Page title
page_id
int
MediaWiki page ID
revision_id
int
Revision ID of the exported version
timestamp
string
Last edit timestamp (ISO 8601)… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/soyjak-wiki.wikipedia_multiple_choice_qa
Galician and Portuguese Multiple-Choice QA Instruction Subsets
Dataset description
This dataset contains two instruction-tuning subsets for multiple-choice question answering in Galician and Portuguese:
gl_wikipedia_multiple_choice_qa (1,486 instances)
pt_wikipedia_multiple_choice_qa (547 instances)
Both subsets are reformatted versions of QA data originally included in the cpt_instruction_datasets collection, adapted here as standalone instruction-style datasets.
Each… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/wikipedia_multiple_choice_qa.qfs-smollm2-135m-wikitext2-gptq-g32-v1
HF workflow 32c6ab05b0ceab1cecdceda838846388
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g32.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g32-v1.wiki2023_plus
Overview
The dataset includes the description from Wikipedia and categories of films published in 2023.
This dataset is used to evaluate the ability of LLM to memorize and extract information described in the document.
See "Where is the Answer? An Empirical Study of Positional Bias for Parametric Knowledge Extraction in Language Model (NAACL2025 Long paper)" for how we use this dataset for training and evaluation.
Data Split
film_doc_all.jsonl includes lines of… See the full description on the dataset page: https://huggingface.co/datasets/omron-sinicx/wiki2023_plus.wikiSQL-ru-datasetturkish-wikipedia-dataset-clean
Turkish Wikipedia Dataset
A cleaned and structured Turkish Wikipedia dataset designed for Turkish language model pretraining, continued pretraining, research, and NLP experiments.
The dataset consists of articles collected from the Turkish Wikipedia (tr.wikipedia.org) and processed into a machine-readable format while preserving important source metadata.
Dataset Summary
Language: Turkish (tr)
Source: Turkish Wikipedia
Domain: General knowledge / encyclopedia… See the full description on the dataset page: https://huggingface.co/datasets/kaan39/turkish-wikipedia-dataset-clean.Wiki_Faiss_Indexes
dataset_info:
features:
- name: text
dtype: string
- name: embeddings
dtype: float32
shape: [384]
configs:
- config_name: default
data_files: "*.parquet"
Wikipedia IVF-OPQ-PQ Vector Database (GPU-Optimized)
A high-performance, GPU-accelerated FAISS vector database built from Wikipedia articles with pre-computed embeddings. This dataset contains approximately 35 million Wikipedia articles with 384-dimensional embeddings using the all-MiniLM-L6-v2… See the full description on the dataset page: https://huggingface.co/datasets/Ram-G/Wiki_Faiss_Indexes.qfs-smollm2-135m-wikitext2-gptq-g64-v1
HF workflow 73f0a12a901c7368794a3a886f55b675
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g64.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g64-v1.
