datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3D-dungeon-crawler-video-v2
3D Dungeon Crawler Video v2
32,000 deterministic 28-second observational Unity episodes.
The canonical split contains 16,000 pretrain, 14,000 training,
1,000 test, and 1,000 evaluation episodes.
Unity renders at 512x288 for supersampling. Videos are stored at
256x144, 30 fps, H.264. Training samples every third frame,
yielding 280 frames and an 18x32 visual-token grid per episode.
manifest.jsonl is authoritative for asset paths. Each record points to one MP4 and one
NPZ… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-video-v2.3D-dungeon-crawler-stratified-continuations-v1
Stratified conditional-continuation benchmark (v1)
Complete and frozen.
A frozen evaluation cohort of 4,000 matched pairs (8,000 episodes)
for estimating and comparing
P(M,Y,R | do(X=1), A=1) P(R=1 | do(X=1), A=1)
from a common set of pre-X contexts shared by every model arm and by the
simulator reference. No arm may select a different prefix set.
manifest.jsonl is immutable and audit.json records the result of every
audit required by issue #53, including the exact-quota… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-stratified-continuations-v1.v2-crawler3D-dungeon-crawler-video-v2-leaky-xor-supplementai-crawler-index
AI Crawler Index
150 web crawlers and AI user agents from 74 operators — what each one is for,
what blocking it costs you, and the IP ranges its operator publishes.
Plus a compiled user-agent regex and the union of 1997 IPv4 and 1062 IPv6
prefixes from 15 operator-published range files.
Home: https://www.pathwren.workers.dev/c/huggingface-datasets/ · CC0 · no signup, no key.
What this is, plainly
This is an independent, non-commercial automated project. It is run… See the full description on the dataset page: https://huggingface.co/datasets/pathwren/ai-crawler-index.3D-dungeon-crawler-video
3D Dungeon Crawler Video
Deterministic 23-second Unity episodes for observational world-model training. The 10,000 episodes are split into 8,000 train, 1,000 validation, and 1,000 test episodes.
Unity renders at 512x288 for supersampling; each clips/*.mp4 is area-downscaled and stored at 256x144 and 30 fps. Training samples every third frame, yielding 230 frames and an 18x32 tokenizer grid. Matching arrays/*.npz files contain dag_ticks, action_tokens, and final_state.
All… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-video.newsxlm-mhtml
MHTML Sources for NewsXLM dataset
Columns Explanation
mhtml: MHTML snapshots of pages where wtl-uid and wtl-parent-uid attributes have been added to every element in <body>, following the WTL algorithm.
v1.3-final-crawlerv3-crawlernewsxlm
NewsXLM Dataset
The first large-scale multilingual dataset for news web page attribute extraction, containing 29,081 annotated pages from 759 websites across 56 languages.
📊 Per-language Stats
🗃️ MHTML Sources
Columns Explanation
html: Cleaned original HTML
html_en: html with text nodes translated into English
html_wtl: html with wtl-uid and wtl-parent-uid attributes added to each element in <body>, following the WTL algorithmannotations: List of labeled page nodes in the… See the full description on the dataset page: https://huggingface.co/datasets/ispras-crawlers/newsxlm.sft-crawler-backup3D-dungeon-crawler-interventions
3D Dungeon Crawler Interventions
9,000 deterministic 23-second Unity intervention episodes: 500 examples for each do(variable=0/1) arm over D, Z, V, W, X, M, A, Y, and R. Every do(0)/do(1) pair shares its layout seed and aligned per-variable random stream. Pairs stay in the same split; every arm contributes 400 train, 50 validation, and 50 test episodes.
Unity renders at 512x288 for supersampling; each clips/*.mp4 is area-downscaled and stored at 256x144 and 30 fps. Training… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-interventions.GUI-Net-Crawler
How to use this data?
After download this repo, use cat to get zip file:
cat baidu_wiki_part_* > merge.zip
Then simply, unzip this zip file
unzip merge.zip
What is in this data?
Image(Screenshot)
Raw images are in images folder.
/wikihow$ ls data/images | head -5
1111-4.jpg
111-15.jpg
1-draw-7.png
20200613_130717.jpg
22-19.jpg
Index page
Index page is a collection of web urls. This is how we start to crawl these websites.
wikihow$ cat… See the full description on the dataset page: https://huggingface.co/datasets/Bofeee5675/GUI-Net-Crawler.vi_knowledge-crawlers
TODOs
Format lại vietceteravihtml.jsonl.xz thành pure text
Crawl bản tiếng Anh
Chuyển vào zin/21-vicnen_alignment_21-of-32G/
Lọc thivien.net.jsonl.xz chỉ giữ lại en, cn, vi
Chuyển vào zinz/50-vicnen_advanced-alignment/
fr-crawler-private
Dataset Card for "fr-crawler-private-mlm"
More Information needed
CRAwLeR-PL
CRAwLeR-PL — Cross-Reference Aware Legal Retrieval (Polish)
CRAwLeR measures context-aware (contextual) chunk retrieval: specifically, legal cross-reference retrieval. To rank the right chunk, a retriever must use information from other chunks that the target chunk cross-references. This is the Polish instance, built on Polish legal Acts obtained through the ELI API.
It is the first dataset for context-aware chunk retrieval to carefully consider construct validity and inspect… See the full description on the dataset page: https://huggingface.co/datasets/macja/CRAwLeR-PL.CRAwLeR-DK
CRAwLeR-DK — Cross-Reference Aware Legal Retrieval (Danish)
CRAwLeR measures context-aware (contextual) chunk retrieval: specifically, legal cross-reference retrieval. To rank the right chunk, a retriever must use information from other chunks that the target chunk cross-references. This is the Danish instance, built on Danish legal documents from Retsinformation.
It is the first dataset for context-aware chunk retrieval to carefully consider construct validity and inspect… See the full description on the dataset page: https://huggingface.co/datasets/macja/CRAwLeR-DK.shelfglance-ai-crawler-blocking
133 of 10,099 Shopify stores block an AI crawler. 6 block the one ChatGPT shops with.
The robots.txt of 10,099 Shopify storefronts, read for twelve AI crawlers. 133 block at least one, almost all of them training crawlers by a copied whole-site rule; 13 block a crawler that fetches at answer time, and no store blocks the shopping crawler while allowing the training one.
One row per Shopify storefront, out of 10,099 read between 29 August 2026 and 2 September 2026, whose… See the full description on the dataset page: https://huggingface.co/datasets/bpm2007/shelfglance-ai-crawler-blocking.ai-crawler-user-agents
AI Crawler User Agents
Machine-readable list of all 28 known AI crawler and agent user-agent strings —
GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Bytespider,
Applebot-Extended, and more — with each bot's operator, purpose, observed
robots.txt compliance, and documented crawl-delay support.
Fields
Field
Description
userAgent
Token to match in robots.txt / server logs (e.g. GPTBot)
operator
Company running the crawler
purpose… See the full description on the dataset page: https://huggingface.co/datasets/osamamumtaz01/ai-crawler-user-agents.douban_crawlerfr_crawler2crawlerlm-html-to-json
CrawlerLM: HTML to JSON Extraction
A synthetic instruction-tuning dataset for training language models to extract structured JSON from HTML.
Dataset Description
This dataset contains HTML paired with structured JSON extraction tasks in chat format. It's designed for fine-tuning small language models to perform structured data extraction from messy, real-world HTML across multiple domains.
Key Features
447 examples in instruction-tuning chat format
Real HTML… See the full description on the dataset page: https://huggingface.co/datasets/espsluar/crawlerlm-html-to-json.crawler-signal-reference
Crawler Signal Reference Library
A 30-document reference library covering every HTML and HTTP signal a web crawler reads from a site. Authored by Joseph W. Anady, founder of ThatDevPro and ThatDeveloperGuy.
Coverage
HTML signals (18 documents)
App + PWA metadata, canonical + alternate, favicon + icons
All meta tags (author, charset, color-scheme, content-language, copyright, generator, keywords, referrer, refresh, robots, theme-color, viewport)
Open Graph… See the full description on the dataset page: https://huggingface.co/datasets/ThatDeveloperGuy13/crawler-signal-reference.fr-crawler-private-mlm
Dataset Card for "fr-crawler-private-mlm"
More Information needed
fr_crawler_mlm
Dataset Card for "fr_crawler"
More Information needed
fr_crawler_class
Dataset Card for "fr_crawler2"
More Information needed
ai-crawlers-ip-and-user-agent-registry
Known AI Crawlers, User-Agents & Verified IP Ranges Registry (2026)
Official aggregated registry of all major AI Scrapers, Search Bots, and LLM Crawlers operating on the web. It includes verified User-Agent strings, policy block targets (for robots.txt), and official CIDR IP ranges for firewall configuration.
Published by Pixel Office EU.
Purpose & Utility
As AI web indexing scales exponentially, website administrators and system engineers face the challenge of… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/ai-crawlers-ip-and-user-agent-registry.zalo-crawler-v17-explanation
Dataset Card for "zalo-crawler-v17-explanation"
More Information needed
sitemap-to-url-crawler-sample-data
Sitemap to URL Crawler
nstantly extract all public URLs from any website's sitemap.xml recursively. Handles nested sitemap indexes automatically. The fastest & cheapest way to build URL lists for RAG pipelines, LLM training, and SEO audits. Zero-config & blazing fast.
What the actor scrapes
Sitemap to URL Crawler — RAG & AI Data Feeder Extract every public URL from any website's sitemap.xml — recursively, instantly, and at scale. Handles nested sitemap indexes… See the full description on the dataset page: https://huggingface.co/datasets/logiover/sitemap-to-url-crawler-sample-data.crawler_v0
