datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webshop-dataessential-web-v1.0
🌐 Essential-Web: Complete 24-Trillion Token Dataset
🏆 Website | 🖥️ Code | 📖 Paper | ☁️ AWS
📋 Dataset Description
Essential-Web is a 24-trillion-token web dataset with document-level metadata designed for flexible dataset curation. The dataset provides metadata including subject matter classification, web page type, content complexity, and document quality scores for each of the 23.6 billion documents.
Researchers can filter and curate specialized datasets… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/essential-web-v1.0.WebVidCorpus-200B
WebOrganizer/Corpus-200B
[Paper] [Website] [GitHub]
This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned with
(1) RefinedWeb filters and
(2) BFF deduplication.
We provide the resulting 200B token corpus annotated with two quality scores, WebOrganizer domains, and k-means scores.
Download the dataset by cloning the repository with Git LFS instead of HuggingFace's load_dataset().
The dataset has the following folder structure:… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/Corpus-200B.weblinx-browsergym
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Xing Han Lù*, Zdeněk Kasner*, Siva Reddy
💾Code
📄Paper
🌐Website
📓Colab
🤖Models💻Explorer
🐦Tweets
🏆Leaderboard
Your browser does not support the video tag.
This dataset was specifically created to allow WebLINX to be used inside the BrowserGym and Agentlab ecosystem. Please see the browsergym repository for more information.
[!NOTE]
The version associated with this library is WebLINX… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/weblinx-browsergym.open-web-math
Keiran Paster*, Marco Dos Santos*, Zhangir Azerbayev, Jimmy Ba
GitHub | ArXiv
| PDF
OpenWebMath is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of 6.3 million documents containing a total of 14.7B tokens. OpenWebMath is intended for use in pretraining and finetuninglarge language models.
You can download the dataset using Hugging Face:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/open-web-math/open-web-math.WebLINX-full
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead:
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Xing Han Lù*, Zdeněk Kasner*, Siva Reddy
💾Code
📄Paper
🌐Website
📓Colab
🤖Models
💻Explorer
🐦Tweets
🏆Leaderboard
Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.webtoepub-libraryWebSight
Dataset Card for WebSight
Dataset Description
WebSight is a large synthetic dataset containing HTML/CSS codes representing synthetically generated English websites, each accompanied by a corresponding screenshot.
This dataset serves as a valuable resource for tasks such as generating UI codes from a screenshot.
It comes in two versions:
v0.1: Websites are coded with HTML + CSS. They do not include real images.
v0.2: Websites are coded with HTML + Tailwind CSS. They do… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/WebSight.TreeOfLife-10M-WEBP
Dataset Card for TreeOfLife-10M-WEBP
Dataset Description
This is an optimized version of the TreeOfLife-10M dataset,
containing over 10 million images covering 454 thousand taxa in the tree of life.
This version has been processed to improve usability and reduce storage requirements while maintaining full compatibility with the original dataset structure.
Dataset Summary
This version modifies the original dataset as follows:
Corrupted files were… See the full description on the dataset page: https://huggingface.co/datasets/birder-project/TreeOfLife-10M-WEBP.web_questions
Dataset Card for "web_questions"
Dataset Summary
This dataset consists of 6,642 question/answer pairs.
The questions are supposed to be answerable by Freebase, a large knowledge graph.
The questions are mostly centered around a single named entity.
The questions are popular ones asked on the web (at least in 2013).
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/web_questions.webvid-10Mbiomedica_webdataset_24M
Dataset Card for Dataset Name
Arxiv: Arxiv
|
Website: Biomedica
|
Training instructions: OpenCLIP
|
Tutorial: Google Colab
BIOMEDICA Dataset is a large-scale, deep-learning-ready biomedical dataset containing over 24M imagecaption pairs and 30M image-references from 6M unique open-source articles. Each… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/biomedica_webdataset_24M.essential-web-1t-sample-fdc-partitioned
🌐 Essential-Web: FDC Level-2 Partitioned Dataset
📋 Dataset Description
This dataset contains a 1 trillion token sample from Essential-Web, partitioned by Free Decimal Correspondence (FDC) level-2 categories. Essential-Web is a 24-trillion-token web dataset with extensive document-level metadata designed to enable rapid dataset curation through SQL-like filtering.
🔍 Free Decimal Correspondence (FDC)
The FDC taxonomy is an open classification system… See the full description on the dataset page: https://huggingface.co/datasets/Research-EAI/essential-web-1t-sample-fdc-partitioned.WEB-Dataset
WorldEngine Bimanual Dataset for Post-training
A large-scale, language-annotated real-robot bimanual manipulation dataset for
post-training robotics foundation models. It spans 90 everyday manipulation tasks
collected with a bimanual YAM follower arm teleoperated by a GELLO leader,
recording joint state, action, and three synchronized camera streams at 60 Hz.
Shared lineage, different story. This dataset shares its hardware, teleoperation
setup, and recording pipeline with the… See the full description on the dataset page: https://huggingface.co/datasets/WorldEngineAI/WEB-Dataset.safebooru-webp-4Mpixel
Safebooru 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/safebooru_full. And all the resized images are maintained here.
There are 5756655 images in total. The maximum ID of these images is 5974383. Last updated at 2025-08-06 08:31:53 JST.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/safebooru-webp-4Mpixel.webagentswebfaq-retrievalWebFAQ Retrieval Dataset
Overview |
Details |
Structure |
Examples |
Considerations |
License |
Citation |
Contact |
Acknowledgement
Overview
The WebFAQ Retrieval Dataset is a carefully filtered and curated subset of the broader WebFAQ Q&A Dataset.It is purpose-built for Information Retrieval (IR) tasks, such as training and evaluating dense or sparse retrieval models in multiple languages.
Each of the… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval.swiss-caselaw-web-ui
Swiss Case Law Open Dataset
962,724 published decisions from Swiss federal, cantonal, and regulatory bodies.
Full text, structured metadata, and daily updates. The March 20, 2026 snapshot contains German, French, and Italian decisions; the export schema also reserves rm for Romansh.
What this is
A structured, searchable archive of Swiss court decisions — from the Federal Supreme Court (BGer) down to cantonal courts in all 26 cantons. Every decision includes the full… See the full description on the dataset page: https://huggingface.co/datasets/ArneH/swiss-caselaw-web-ui.conceptual-captions-12m-webdatasetEdge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.WebInstructSub
🦣 MAmmoTH2: Scaling Instructions from the Web
Project Page: https://tiger-ai-lab.github.io/MAmmoTH2/
Paper: https://arxiv.org/pdf/2405.03548
Code: https://github.com/TIGER-AI-Lab/MAmmoTH2
WebInstruct (Subset)
This repo contains the partial dataset used in "MAmmoTH2: Scaling Instructions from the Web". This partial data is coming mostly from the forums like stackexchange. This subset contains very high-quality data to boost LLM performance through instruction tuning.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/WebInstructSub.webcode2m_purifiedWebCode2M: A Real-World Dataset for Code Generation from Webpage Designs
Features:
image: the screenshot of the webpage.
bbox: the layout information, i.e., the bounding boxes (Bbox) of all the elements in the webpage, which contains the size, position, and hierarchy information.
text: the webpage code text including HTML/CSS code.
scale: the scale of the screenshot, in the format [width, height].
lang: the main language of the text content displayed on the rendered page (excluding HTML/CSS… See the full description on the dataset page: https://huggingface.co/datasets/xcodemind/webcode2m_purified.web-crawl-2026
Web Crawl 2026
A large-scale web crawl dataset for language model pretraining, collected by the OpenTransformer project.
Dataset Description
This dataset contains text extracted from web pages crawled directly from the internet using custom high-throughput crawlers. All data is freshly scraped.
Data Format
Each record is a JSON line (gzipped) with fields:
text: extracted text content (200-200,000 chars)
url: source URL
domain: source domain
timestamp: crawl… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-2026.conceptual-captions-12m-webdataset-metadata
Conceptual Captions 12M — Webshart metadata indices
Per-shard webshart metadata indices for
laion/conceptual-captions-12m-webdataset:
1,100 JSON files under data/, one per source tar shard, mirroring the source's shard layout.
Each index records every tar member's byte offset and length (enabling ranged reads without
downloading whole shards), image geometry (width/height for aspect bucketing), and — as of
August 2026 — embedded captions for all 10,994,853 samples, coalesced… See the full description on the dataset page: https://huggingface.co/datasets/webshart/conceptual-captions-12m-webdataset-metadata.webfiddle-internet-raw-cache-datasetA dataset of different files that robots tried to crawl through webfiddle.net
Mostly html files but other files too pdfs, images, binary- i have no idea what is in here at this stage - but gives an interesting idea of what crawlers like to visit and could be the basis of interesting SEO or coding LLM reasearch.
Collected as part of my work on web simulators.
https://webfiddle.net JS/CSS editor for the web, https://websim.netwrck.com Coding Editor for the web.
https://x.com/leeleepenkman
Its… See the full description on the dataset page: https://huggingface.co/datasets/lee101/webfiddle-internet-raw-cache-dataset.webgym_tasks
WebGym Tasks Dataset
Dataset Description
This dataset contains web navigation tasks for training and evaluating autonomous web agents. Each task consists of a natural language instruction that describes an action to be performed on a specific website, along with evaluation criteria and metadata.
Dataset Summary
Total Training Tasks: 292,092
Total Test Tasks: 1,167
Domains: Multiple domains including Lifestyle & Leisure, Sports & Fitness, and more
Source… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/webgym_tasks.webis-touche2020-v3
Touche2020Retrieval.v3
An MTEB dataset
Massive Text Embedding Benchmark
Touché Task 1: Argument Retrieval for Controversial Questions
Task category
t2t
Domains
Academic
Reference
https://github.com/castorini/touche-error-analysis
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["Touche2020Retrieval.v3"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/webis-touche2020-v3.RoG-webqsp
Dataset Card for "RoG-webqsp"
More Information needed
MegaMath-Web-Pro-Max
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
The Curation of MegaMath-Web-Pro-Max
Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year;
Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data;
Step 3: Training a fasttext carefully with proper preprocessing;
Step 4: Filtering documents with a threshold (i.e., 0.4);
Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.
