datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webchain
WebChain v2
A large-scale, human-annotated dataset of real-world web interaction trajectories for training and evaluating web agents.
[Paper] [Code] [Dataset]
WebChain captures how people complete real tasks on live websites. It is designed for agents that must both identify the correct interface element and reason through a sequence of actions. Each trajectory aligns screenshots, web structure, grounded actions, and reasoning signals instead of treating web navigation as… See the full description on the dataset page: https://huggingface.co/datasets/webagentlab/webchain.pk-map-statsMegaMath-Web-Pro-Max
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
The Curation of MegaMath-Web-Pro-Max
Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year;
Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data;
Step 3: Training a fasttext carefully with proper preprocessing;
Step 4: Filtering documents with a threshold (i.e., 0.4);
Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.osm-polygon-website-tag
OSM Polygon Website Dataset
OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files.
At a glance
Polygons
1,726,474
With extracted text
1,192,980
Words of text
407,685,655
Languages
397
Regional sources
386 / 386
Duplicate objects removed
104,927
Candidates rejected
868,905,743
Status
In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.web-server-logs
Web Server Access Logs (Synthetic) (Free Sample)
This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables.
Realistic HTTP access logs from a simulated SaaS company running an
e-commerce API and marketing website. 50,000 requests across 3 servers
over 12 months.
Includes realistic patterns: weekday/weekend traffic variation, peak hours,
seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and
database outage) for anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/web-server-logs.webgpt_comparisonsWebGPT Comparisons contains all of the comparisons marked as suitable for reward modelling from the WebGPT paper.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.WebTerminal
Terminal/CLI Web Text
A filtered extract of terminal and command-line content from two large web-text corpora, designed for upsampling agentic-adjacent data during pretraining.
Subsets
Subset
Rows
Tokens
Size
Quality
clean (default)
2.33M
4.6B
11 GB
~98% terminal content
unfiltered
61.3M
359B
962 GB
~15% terminal content
from datasets import load_dataset
# Load the clean subset (default)
ds = load_dataset("AdaMLLab/WebTerminal")
# Load the unfiltered… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/WebTerminal.KOREAN-WEBTEXT
KOREAN-WEBTEXT
KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources:
cc100
oscar-corpus/OSCAR-2201
oscar-corpus/OSCAR-2109
oscar-corpus/OSCAR-2301
ontocord/CulturaY
Additional credible internet sources collected by out team
(We are working to add more sources)
The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.general-web-it-202608
General · Web · Italian · 2026-08
Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
21,065,052 documents and 66,158,573,443 characters of Italian prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-ita_Latn
21,065,052
66,158,573,443
epfml/FineWeb2-HQ, ita_Latn
The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.common-corpus-sample-open-webgeneral-web-fr-202608
General · Web · French · 2026-08
French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
31,999,309 documents and 118,346,333,763 characters of French prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-fra_Latn
31,999,309
118,346,333,763
epfml/FineWeb2-HQ, fra_Latn
The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.pretrain-web-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
8,689,580,607 (8.7B)
Trainable tokens
8,689,580,607 (8.7B)
Documents
281,846
Shards
89
UTF-8 bytes
37,540,769,483
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix-long-context.webis-touche2020-decontaminated
webis-touche2020 (Decontaminated)
A decontaminated version of the webis-touche2020 dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed.
Decontamination methodology
Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files):
Pass 1: Exact hash matching
All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/webis-touche2020-decontaminated.webis-touche2020-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/webis-touche2020-qrels.WebArena-Verified
WebArena-Verified
Dataset description
WebArena-Verified is a curated benchmark dataset of web tasks designed for reproducible
evaluation of web agents across multiple realistic websites.
Sources
GitHub repository: webarena-verified
Original WebArena benchmark: webarena.dev
Splits
full: 812 rows
hard: 258 rows
Tasks per site
Counts below are task counts grouped by category. Tasks with more than one site are grouped
under multi-category… See the full description on the dataset page: https://huggingface.co/datasets/AmineHA/WebArena-Verified.YFCC14M_subset_webdatasetSurvivor
📚 FinePDFs-Edu
350B+ of highly educational tokens from PDFs 📄
What is it?
📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages.
FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset.
We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/Web3Survivor/Survivor.megamath-web-prowebgpu-bench-leaderboardso101_main_bin_2cameras_webThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 53,
"total_frames": 9844,
"total_tasks": 1,
"total_videos": 106,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:53"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/guanfengliu/so101_main_bin_2cameras_web.seeclick-web-commercial-mlx
SeeClick Web Commercial Dataset (MLX-VLM Format)
Commercial-use friendly GUI grounding dataset from SeeClick Web data.
Apache 2.0 licensed - safe for commercial applications.
Dataset Description
This dataset contains ~20k examples for training Vision-Language Models to predict
click coordinates given a screenshot and instruction. Derived from SeeClick Web
crawled data (Apache 2.0).
Key Features
License: Apache 2.0 (commercial use allowed)
Format: MLX-VLM… See the full description on the dataset page: https://huggingface.co/datasets/pierretokns/seeclick-web-commercial-mlx.cosmopedia_web_textbooks_logprobsduplo-pick-and-place-yellow-640x480-3-camerasThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/webstep/duplo-pick-and-place-yellow-640x480-3-cameras.dead-web-commoncrawl
Dead-Web Common Crawl — a longitudinal host-reachability panel (2018–2026)
· Hugging Face
· Kaggle
· License: CC BY 4.0
An open dataset labeling the reachability of 172,959,928 registered domains across 80 monthly
Common Crawl archives (CC-MAIN-2018-05 → CC-MAIN-2026-21), built from Common Crawl's robotstxt
subset and calibrated against a live re-probe to separate genuinely-dead domains from those merely
blocking crawlers or dropped from the… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/dead-web-commoncrawl.TopicAnnotations-Llama-3.1-8B
WebOrganizer/TopicAnnotations-Llama-3.1-8B
[Paper] [Website] [GitHub]
This dataset contains 1M web pages annotated with topic labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/TopicClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page
top_choice_index: Index of the most… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-8B.clean-visual-webarena-classifiedswebshop_instructionsWebDocumentDescriptors
Data release for the paper Task-Agnostic Web Document Annotation with LLM-Generated Descriptors (forthcoming).
The descriptors are generated via a task-agnostic data annotation pipeline described in the paper (link coming soon).
This Hugging Face dataset repository contains 5 distinct datasets:
a descriptor-annotated version of a 10 billion token (~15 million document) sample of FineWeb.
The 800k label descriptor schema
The 500k document sample of FineWeb used to develop the schema along… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/WebDocumentDescriptors.rank-distillmThis dataset contains the training run files from the paper Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-ranking for training queries from MS MARCO passage re-ranked by RankZephyr, a large monoELECTRA model or a large Set-Encoder model. These run files can be used to distill smaller and more efficient models while upholding effectiveness.
The files __colbert__msmarco-passage-train-judged.parquet and __bm25__msmarco-passage-train-judged.parquet… See the full description on the dataset page: https://huggingface.co/datasets/webis/rank-distillm.
