datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia_en
wikipedia_en
This is a curated Wikipedia English dataset for use with the II-Commons project.
Dataset Details
Dataset Description
This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1004
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.wikipedia-pageviews
Wikipedia Article Pageviews
This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia.
It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago.
The fetcher runs in a scheduled GitHub Actions workflow, which is available here.
The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1008
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008.WikiRAG-TR
Dataset Summary
WikiRAG-TR is a dataset of 6K (5999) question and answer pairs which synthetically created from introduction part of Turkish Wikipedia Articles. The dataset is created to be used for Turkish Retrieval-Augmented Generation (RAG) tasks.
Dataset Information
Number of Instances: 5999 (5725 synthetically generated question-answer pairs, 274 augmented negative samples)
Dataset Size: 20.5 MB
Language: Turkish
Dataset License: apache-2.0
Dataset Category:… See the full description on the dataset page: https://huggingface.co/datasets/Metin/WikiRAG-TR.WikiProfile
WikiProfile
WikiProfile is a factual knowledge benchmark for evaluating how well language models encode and recall factual knowledge. It comprises 2,150 facts, each paired with 10 questions, for a total of 21,500 question instances.
Each fact is grounded in the first paragraph (summary) of an English Wikipedia page and is defined as a proposition between two entities, a subject and an object (e.g., "Oasis played their first gig at the Boardwalk club" → subject: Oasis, object:… See the full description on the dataset page: https://huggingface.co/datasets/google/WikiProfile.ADFA_MappingLDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1005
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005.wikidata1
Wikipedia N-Link Basins
A novel graph-theoretic analysis of Wikipedia's internal link structure, revealing deterministic "basins of attraction" under N-link traversal rules.
Dataset Description
This dataset demonstrates that Wikipedia's 17.9 million pages partition into coherent regions when following a simple rule: from any page, always follow the Nth link. Every page eventually reaches a cycle, and pages sharing the same terminal cycle form a basin of attraction.… See the full description on the dataset page: https://huggingface.co/datasets/mgmacleod/wikidata1.wikipedia_quality_wikirankDatasets with WikiRank quality score as of 1 August 2024 for 47 million Wikipedia articles in 55 language versions (also simplified version for each language in separate files).
The WikiRank quality score is a metric designed to assess the overall quality of a Wikipedia article. Although its specific algorithm can vary depending on the implementation, the score typically combines several key features of the Wikipedia article.
Why It’s Important
Enhances Trust: For readers and… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/wikipedia_quality_wikirank.cosmopedia-wikihow-chunked
Overview
This dataset is a chunked version of a subset of data in the Cosmopedia dataset curated by Hugging Face.
Specifically, we have only used a subset of Wikihow articles from the Cosmopedia dataset, and each article has been split into chunks containing no more than 2 paragraphs.
Dataset Structure
Each record in the dataset represents a chunk of a larger article, and contains the following fields:
doc_id: A unique identifier for the parent article
chunk_id: A unique… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/cosmopedia-wikihow-chunked.polaris-wikitables-v2
WikiTables v2
WikiTables is one of six datasets in Polaris: Learning to Generate Table Descriptions from
Retrieval Feedback, alongside aw, arctic, lter, ecir,
and wtr.
It holds 3,361 tables scraped from Wikipedia articles and 57 keyword queries over them. For each
query–table pair, a person scored how well that table answers that query; those scores are the
relevance judgments, and they live in qrels.csv.
The tables have no names. What describes a table is its column names and… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-wikitables-v2.GPT-wiki-intro
GPT Wiki Intro
Overview
Dataset for training models to classify human written vs GPT/ChatGPT generated text.
This dataset contains Wikipedia introductions and GPT (Curie) generated introductions for 150k topics.
Prompt used for generating text
200 word wikipedia style introduction on '{title}'
{starter_text}
where title is the title for the wikipedia page, and starter_text is the first seven words of the wikipedia introduction.
Here's an example of prompt used to… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.wikimovies
Wikipedia Movies Dataset
Dataset Description
This dataset contains 58.1k movie information scraped from Wikipedia's "List of films" pages. The data includes basic movie metadata, infobox information, and introductory text from individual movie Wikipedia pages.
NOTE: This is an uncleaned dataset containing raw scraped data. The content is sourced from Wikipedia and is not owned by the dataset creator. All content remains under Wikipedia's licensing terms.… See the full description on the dataset page: https://huggingface.co/datasets/yashassnadig/wikimovies.bergson-wikitext-gpt2-leaderboard-bank
bergson leaderboard: retrain banks, scores and LDS/QLD results (WikiText GPT-2)
Everything behind the numbers on the bergson leaderboard,
for the model at EleutherAI/bergson-wikitext-gpt2-leaderboard.
path
what it is
bank/
the LDS ground truth: 100 random leave-1%-out subsets of the 4,608 training chunks (subsets.json) and each subset's measured loss change on the 50 test queries (validation.csv)
random/retrained/{base,subset_0..99}
the retrained models themselves… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/bergson-wikitext-gpt2-leaderboard-bank.simple_english_wikipediaOnly for the reaseaching usage.
The original data from http://sbert.net/datasets/simplewiki-2020-11-01.jsonl.gz.
We use nq_distilbert-base-v1 model encode all the data to the PyTorch Tensors. And normalize the embeddings by using sentence_transformers.util.normalize_embeddings.
How to use
See notebook Wikipedia Q&A Retrieval-Semantic Search
Installing the package
!pip install sentence-transformers==2.3.1
The converting process
# the whole process takes… See the full description on the dataset page: https://huggingface.co/datasets/aisuko/simple_english_wikipedia.wikipedia-image-requests
Daily image-request counts for Wikipedia article images
Daily counts of how often each of ~1.19 million Wikipedia article images was requested from
Wikimedia's image servers, attributed to the wiki whose page the request came from, together with
the article each image appears on.
406,263,728 daily observations covering 1,186,150 image files across 534,067 articles in
nine Wikipedias, from 2025-09-01 to 2026-09-05 (370 days).
Every file in the set appears on exactly one… See the full description on the dataset page: https://huggingface.co/datasets/lgelauff/wikipedia-image-requests.Wiki_Live_Challenge
Wiki Live Challenge Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems.
Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.polaris-wikitables-v1
WikiTables v1
WikiTables is one of six datasets in Polaris: Learning to Generate Table Descriptions from
Retrieval Feedback, alongside aw, arctic, lter, ecir,
and wtr.
It holds 3,361 tables scraped from Wikipedia articles and 57 keyword queries over them. For each
query–table pair, a person scored how well that table answers that query; those scores are the
relevance judgments, and they live in qrels.csv.
The tables have no names. What describes a table is its column names and… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-wikitables-v1.wikipedia-citation-indexDataset with citation indexes as of 1 August 2024 for 47 million Wikipedia articles in 55 language versions. Research: ArXiv
wikivoyage-eu-city-embeddings
Dataset Card for Dataset Name
This dataset comprises abstracts from Wikivoyage for 160 European cities along with their corresponding country names, coordinates, and populations. The embeddings are derived from the GTE-Large model, incorporating data from the city, country, population, and abstract columns.
Dataset Sources
Wikivoyage data
World cities database
moonlighter-2-wiki-data
Moonlighter 2 relic prices and weapons dataset
An open, versioned export from www.moonlighter2.wiki, an independent and unofficial Moonlighter 2 wiki published as The Endless Ledger.
This release contains two small, research-friendly tables:
Relic prices: 159 published relics, including 154 rows with sourced prices and 5 rows whose unknown prices are intentionally left blank.
Weapons: 30 published weapons with weapon type, effects, upgrade information, verification state and… See the full description on the dataset page: https://huggingface.co/datasets/zyzw10086/moonlighter-2-wiki-data.thai-wikiqa-tha-qaretrievalref: https://aiforthai.in.th/
small-GPT-wiki-intro-features
Small-GPT-wiki-intro-features dataset
This dataset is based on aadityaubhat/GPT-wiki-intro.
It contains 100k randomly selected texts (50k from Wikipedia and 50k generated by ChatGPT).
For each text, various complexity measures were calculated, including e.g. readibility, lexical richness etc.
It can be used for text classification or analysis of linguistic features of human-generated and ChatGPT-generated texts.
Dataset structure
Features were calculated using… See the full description on the dataset page: https://huggingface.co/datasets/julia-lukasiewicz-pater/small-GPT-wiki-intro-features.p2c_polite_wikiwikipedia-22-12-en-voyage-embedsimple_english_wikipedia_p0Only for the researching usage.
The converting process below.
# Setting the env
os.environ['DATASET_URL']='http://sbert.net/datasets/simplewiki-2020-11-01.jsonl.gz'
os.environ['MODEL_NAME']='multi-qa-MiniLM-L6-cos-v1'
# Loading the dataset
import json
import gzip
from sentence_transformers.util import http_get
http_get(os.getenv('DATASET_URL'), os.getenv('DATASET_NAME'))
passages=[]
with gzip.open(os.getenv('DATASET_NAME'), 'rt', encoding='utf8') as fIn:
for line in fIn:… See the full description on the dataset page: https://huggingface.co/datasets/aisuko/simple_english_wikipedia_p0.wiki-stance-en
