datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WebTerminal
Terminal/CLI Web Text
A filtered extract of terminal and command-line content from two large web-text corpora, designed for upsampling agentic-adjacent data during pretraining.
Subsets
Subset
Rows
Tokens
Size
Quality
clean (default)
2.33M
4.6B
11 GB
~98% terminal content
unfiltered
61.3M
359B
962 GB
~15% terminal content
from datasets import load_dataset
# Load the clean subset (default)
ds = load_dataset("AdaMLLab/WebTerminal")
# Load the unfiltered… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/WebTerminal.general-web-it-202608
General · Web · Italian · 2026-08
Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
21,065,052 documents and 66,158,573,443 characters of Italian prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-ita_Latn
21,065,052
66,158,573,443
epfml/FineWeb2-HQ, ita_Latn
The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.general-web-fr-202608
General · Web · French · 2026-08
French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
31,999,309 documents and 118,346,333,763 characters of French prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-fra_Latn
31,999,309
118,346,333,763
epfml/FineWeb2-HQ, fra_Latn
The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.Survivor
📚 FinePDFs-Edu
350B+ of highly educational tokens from PDFs 📄
What is it?
📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages.
FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset.
We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/Web3Survivor/Survivor.exp-pool-olmo-web-dolma2-tokenized
Locus EXP OLMo Web - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.medium-web-pentesting
Medium Web Pentesting Articles
Dataset Description
A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body.
This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.kurdish-web
Dataset Card for Kurdish Web Corpus (Deduplicated)
Dataset Summary
A multilingual corpus of web text in three Kurdish/Zaza varieties — Kurmanji Kurdish
(kmr_Latn), Sorani Kurdish (ckb_Arab), and Zazaki (diq_Latn) — scraped from Kurdish-language
websites, language-identified with GlotLID, and
deduplicated (exact + MinHash near-duplicate removal). See the "Dataset Creation" section below
for the full pipeline.
Rows, by language config:
config
language
script… See the full description on the dataset page: https://huggingface.co/datasets/muzaffercky/kurdish-web.WEBPRMBENCH
WebPRMBench
The first comprehensive evaluation benchmark for Web Process Reward Models
Published at ICLR 2026
Paper | Code | Website | Collection | Demo
Overview
WebPRMBench is the first comprehensive evaluation benchmark dedicated to Web Process Reward Models (WebPRMs). It evaluates how well a reward model can judge the quality of web agent actions during long-horizon web navigation. Each instance presents a web state (page context, trajectory history, user… See the full description on the dataset page: https://huggingface.co/datasets/ZYao720/WEBPRMBENCH.unclean-web
🕸️ Unclean Web
Raw, unfiltered web data scraped across a wide variety of sites and packaged
for language model pre-training, fine-tuning, and research.
📊 Dataset Statistics
Metric
Value
Total Pages
18,896
Total Token Estimate
24.38M
Unique Sources
54
Schema Version
3.0
Last Updated
2026-06-07 00:36 UTC
🗂️ Available Splits / Subsets
Split
Description
Format
full (per batch)
Complete raw scrape, all columns… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/unclean-web.webui
🌐 WebUI
A high-quality dataset for training AI models to understand and generate modern websites.
The dataset contains structured webpage samples including HTML, screenshots, UI element annotations, semantic labels, color palettes, fonts, layout metadata, and accessibility information. Each sample represents a complete webpage that can be used for web generation, UI understanding, or multimodal training. :contentReference[oaicite:0]{index=0}
✨ Features
📄… See the full description on the dataset page: https://huggingface.co/datasets/xlelords/webui.mogan-turkish-web-long
moganai/mogan-turkish-web long filtered Turkish texts
Source: moganai/mogan-turkish-web (config: default, revision: d773a0efd1b7daf72c7909c83dbd385e7e3564d7).
Rows contain 4,000–16,000 characters and passed the
iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft
information-density filters. Selected rows: 2,805,010.
Generated by process_hf_dataset.py. See summary.json for counts and thresholds.
tiny-webtext
Tiny WebText
The Tiny WebText dataset is designed to help models learn about perception on web text while neutralizing the bias of the source text using critical thinking methods. By providing a rich and diverse set of texts, I aim to improve the ability of models to understand and analyze information in a more objective and unbiased manner.
This dataset can be used to train and evaluate natural language processing and machine learning models, with the goal of improving their… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-webtext.axya-tech-websearch
Dhivehi Combined Dataset
Overview
This dataset combines three 30K Dhivehi language corpora from the Leipzig Corpora Collection into a single unified CSV file containing 90,000 sentences. The dataset provides a comprehensive resource for Dhivehi language processing, combining data from Wikipedia, news sources, and web crawls.
Dataset Composition
The dataset consists of three distinct sources:
Wikipedia (2021): 30,000 sentences from Dhivehi Wikipedia
News… See the full description on the dataset page: https://huggingface.co/datasets/axmeeabdhullo/axya-tech-websearch.diverse-websearch-3.5k
Diverse WebSearch 3.5k
Diverse WebSearch 3.5K is a small web research dataset containing 3,500 diverse web pages collected from search results. Each row contains a source URL, extracted markdown content, a concise summary, and image URLs found on the page.
Dataset Details
This dataset is intended for learning and experimentation with:
webpage summarization
retrieval-augmented generation
search result understanding
document cleaning
synthetic QA generation
dataset… See the full description on the dataset page: https://huggingface.co/datasets/soumikmahato/diverse-websearch-3.5k.spai-ss6-corpus-medical-health-web
SPAI SS6 Thai Medical Health Web Corpus
Thai public medical and health web articles collected by the local scraping pipeline.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: default
Rows in canonical config: 3,660
Parquet size in canonical config: 0.01 GB
Source license: other… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-health-web.spai-ss6-corpus-wangchanlion-web
SPAI SS6 WangchanLION Web Corpus Index
Index repo for the WangchanLION-Web corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: wangchanlion_web
Rows in canonical config: 557,502
Parquet size in canonical config: 1.97 GB
Source license: odc-by… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-wangchanlion-web.role-play-bench
Role-play Benchmark
A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios.
Dataset Summary
Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?". Instead… See the full description on the dataset page: https://huggingface.co/datasets/web3w/role-play-bench.swiss-web-premium-ch
*.ch Swiss Web Premium (A+)
Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- Full provenance -- PII-redacted -- RAG-ready -- SFT-formatted
A production-grade Swiss web corpus from the .ch TLD namespace. 110,491 documents independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Built for LLM training, RAG pipelines, SFT fine-tuning, and multilingual NLP.
OptiTransferData Portfolio
Premium sovereign web corpora for… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch.OpenBookQA-Japanese-masked
OpenBookQA-Japanese-masked
与えられた問題に対して4つの選択肢から答えを選択するデータセット
allenai/openbookqaをcyberagent/calm3-22b-chatで翻訳
5,957件
train split: 4,956件(4,957件の内1件削除)
validation split: 500件
test split: 499件(500件の内1件削除)
ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施
Format
データセットの構成は以下
{
"idx": ID,
"id": 元ID,
"question_stem_en": 英語の質問文,
"choices_en": {
"text": 選択肢の文章,
"label": 選択肢の記号,
}… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/OpenBookQA-Japanese-masked.swiss-web-premium-ch-full
*.ch Swiss Web Premium (A+) -- Full Dataset
Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- 22 production files -- 554.1 MB
The complete production release of OptiTransfer's Swiss web corpus. 110,491 documents from the .ch TLD namespace, independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Delivered in Parquet, JSONL, language splits, and pre-built RAG chunks.
This is the full commercial dataset. For evaluation, see the free… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch-full.
