CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01osazuwa /3D-dungeon-crawler-video-v2 3D Dungeon Crawler Video v2 32,000 deterministic 28-second observational Unity episodes. The canonical split contains 16,000 pretrain, 14,000 training, 1,000 test, and 1,000 evaluation episodes. Unity renders at 512x288 for supersampling. Videos are stored at 256x144, 30 fps, H.264. Training samples every third frame, yielding 280 frames and an 18x32 visual-token grid per episode. manifest.jsonl is authoritative for asset paths. Each record points to one MP4 and one NPZ… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-video-v2.video10K<n<100K1 likes4.4k downloads16h agoHugging Face02osazuwa /3D-dungeon-crawler-stratified-continuations-v1 Stratified conditional-continuation benchmark (v1) Complete and frozen. A frozen evaluation cohort of 4,000 matched pairs (8,000 episodes) for estimating and comparing P(M,Y,R | do(X=1), A=1) P(R=1 | do(X=1), A=1) from a common set of pre-X contexts shared by every model arm and by the simulator reference. No arm may select a different prefix set. manifest.jsonl is immutable and audit.json records the result of every audit required by issue #53, including the exact-quota… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-stratified-continuations-v1.videovideo-classification1K<n<10K0 likes2k downloads4h agoHugging Face03enigmare /v2-crawler1 likes895 downloads25d agoHugging Face04osazuwa /3D-dungeon-crawler-video-v2-leaky-xor-supplementvideon<1K0 likes872 downloads11h agoHugging Face05pathwren /ai-crawler-index AI Crawler Index 150 web crawlers and AI user agents from 74 operators — what each one is for, what blocking it costs you, and the IP ranges its operator publishes. Plus a compiled user-agent regex and the union of 1997 IPv4 and 1062 IPv6 prefixes from 15 operator-published range files. Home: https://www.pathwren.workers.dev/c/huggingface-datasets/ · CC0 · no signup, no key. What this is, plainly This is an independent, non-commercial automated project. It is run… See the full description on the dataset page: https://huggingface.co/datasets/pathwren/ai-crawler-index.tabular1K<n<10K0 likes622 downloads8d agoHugging Face06osazuwa /3D-dungeon-crawler-video 3D Dungeon Crawler Video Deterministic 23-second Unity episodes for observational world-model training. The 10,000 episodes are split into 8,000 train, 1,000 validation, and 1,000 test episodes. Unity renders at 512x288 for supersampling; each clips/*.mp4 is area-downscaled and stored at 256x144 and 30 fps. Training samples every third frame, yielding 230 frames and an 18x32 tokenizer grid. Matching arrays/*.npz files contain dag_ticks, action_tokens, and final_state. All… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-video.video1K<n<10K0 likes392 downloads1mo agoHugging Face07ispras-crawlers /newsxlm-mhtml MHTML Sources for NewsXLM dataset Columns Explanation mhtml: MHTML snapshots of pages where wtl-uid and wtl-parent-uid attributes have been added to every element in <body>, following the WTL algorithm. text10K<n<100K0 likes390 downloads9mo agoHugging Face08enigmare /v1.3-final-crawler0 likes304 downloads1mo agoHugging Face09enigmare /v3-crawler0 likes223 downloads9d agoHugging Face10ispras-crawlers /newsxlm NewsXLM Dataset The first large-scale multilingual dataset for news web page attribute extraction, containing 29,081 annotated pages from 759 websites across 56 languages. 📊 Per-language Stats 🗃️ MHTML Sources Columns Explanation html: Cleaned original HTML html_en: html with text nodes translated into English html_wtl: html with wtl-uid and wtl-parent-uid attributes added to each element in <body>, following the WTL algorithmannotations: List of labeled page nodes in the… See the full description on the dataset page: https://huggingface.co/datasets/ispras-crawlers/newsxlm.imagetoken-classification10K<n<100K2 likes199 downloads9mo agoHugging Face11enigmare /sft-crawler-backup0 likes185 downloads1mo agoHugging Face12osazuwa /3D-dungeon-crawler-interventions 3D Dungeon Crawler Interventions 9,000 deterministic 23-second Unity intervention episodes: 500 examples for each do(variable=0/1) arm over D, Z, V, W, X, M, A, Y, and R. Every do(0)/do(1) pair shares its layout seed and aligned per-variable random stream. Pairs stay in the same split; every arm contributes 400 train, 50 validation, and 50 test episodes. Unity renders at 512x288 for supersampling; each clips/*.mp4 is area-downscaled and stored at 256x144 and 30 fps. Training… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-interventions.video1K<n<10K0 likes133 downloads1mo agoHugging Face13Bofeee5675 /GUI-Net-Crawler How to use this data? After download this repo, use cat to get zip file: cat baidu_wiki_part_* > merge.zip Then simply, unzip this zip file unzip merge.zip What is in this data? Image(Screenshot) Raw images are in images folder. /wikihow$ ls data/images | head -5 1111-4.jpg 111-15.jpg 1-draw-7.png 20200613_130717.jpg 22-19.jpg Index page Index page is a collection of web urls. This is how we start to crawl these websites. wikihow$ cat… See the full description on the dataset page: https://huggingface.co/datasets/Bofeee5675/GUI-Net-Crawler.2 likes91 downloads1y agoHugging Face14tiendung /vi_knowledge-crawlers TODOs Format lại vietceteravihtml.jsonl.xz thành pure text Crawl bản tiếng Anh Chuyển vào zin/21-vicnen_alignment_21-of-32G/ Lọc thivien.net.jsonl.xz chỉ giữ lại en, cn, vi Chuyển vào zinz/50-vicnen_advanced-alignment/ text0 likes69 downloads3y agoHugging Face15factored /fr-crawler-private Dataset Card for "fr-crawler-private-mlm" More Information needed text1K<n<10K0 likes62 downloads3y agoHugging Face16macja /CRAwLeR-PL CRAwLeR-PL — Cross-Reference Aware Legal Retrieval (Polish) CRAwLeR measures context-aware (contextual) chunk retrieval: specifically, legal cross-reference retrieval. To rank the right chunk, a retriever must use information from other chunks that the target chunk cross-references. This is the Polish instance, built on Polish legal Acts obtained through the ELI API. It is the first dataset for context-aware chunk retrieval to carefully consider construct validity and inspect… See the full description on the dataset page: https://huggingface.co/datasets/macja/CRAwLeR-PL.texttext-retrieval10K<n<100K0 likes57 downloads3mo agoHugging Face17macja /CRAwLeR-DK CRAwLeR-DK — Cross-Reference Aware Legal Retrieval (Danish) CRAwLeR measures context-aware (contextual) chunk retrieval: specifically, legal cross-reference retrieval. To rank the right chunk, a retriever must use information from other chunks that the target chunk cross-references. This is the Danish instance, built on Danish legal documents from Retsinformation. It is the first dataset for context-aware chunk retrieval to carefully consider construct validity and inspect… See the full description on the dataset page: https://huggingface.co/datasets/macja/CRAwLeR-DK.texttext-retrieval10K<n<100K0 likes53 downloads3mo agoHugging Face18bpm2007 /shelfglance-ai-crawler-blocking 133 of 10,099 Shopify stores block an AI crawler. 6 block the one ChatGPT shops with. The robots.txt of 10,099 Shopify storefronts, read for twelve AI crawlers. 133 block at least one, almost all of them training crawlers by a copied whole-site rule; 13 block a crawler that fetches at answer time, and no store blocks the shopping crawler while allowing the training one. One row per Shopify storefront, out of 10,099 read between 29 August 2026 and 2 September 2026, whose… See the full description on the dataset page: https://huggingface.co/datasets/bpm2007/shelfglance-ai-crawler-blocking.textn<1K0 likes45 downloads17d agoHugging Face19osamamumtaz01 /ai-crawler-user-agents AI Crawler User Agents Machine-readable list of all 28 known AI crawler and agent user-agent strings — GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Bytespider, Applebot-Extended, and more — with each bot's operator, purpose, observed robots.txt compliance, and documented crawl-delay support. Fields Field Description userAgent Token to match in robots.txt / server logs (e.g. GPTBot) operator Company running the crawler purpose… See the full description on the dataset page: https://huggingface.co/datasets/osamamumtaz01/ai-crawler-user-agents.textn<1K0 likes41 downloads9d agoHugging Face20cloudd-11 /douban_crawlerimage2 likes39 downloads2y agoHugging Face21edanigoben /fr_crawler2tabular1M<n<10M0 likes36 downloads3y agoHugging Face22espsluar /crawlerlm-html-to-json CrawlerLM: HTML to JSON Extraction A synthetic instruction-tuning dataset for training language models to extract structured JSON from HTML. Dataset Description This dataset contains HTML paired with structured JSON extraction tasks in chat format. It's designed for fine-tuning small language models to perform structured data extraction from messy, real-world HTML across multiple domains. Key Features 447 examples in instruction-tuning chat format Real HTML… See the full description on the dataset page: https://huggingface.co/datasets/espsluar/crawlerlm-html-to-json.texttext-generationn<1K2 likes36 downloads9mo agoHugging Face23ThatDeveloperGuy13 /crawler-signal-reference Crawler Signal Reference Library A 30-document reference library covering every HTML and HTTP signal a web crawler reads from a site. Authored by Joseph W. Anady, founder of ThatDevPro and ThatDeveloperGuy. Coverage HTML signals (18 documents) App + PWA metadata, canonical + alternate, favicon + icons All meta tags (author, charset, color-scheme, content-language, copyright, generator, keywords, referrer, refresh, robots, theme-color, viewport) Open Graph… See the full description on the dataset page: https://huggingface.co/datasets/ThatDeveloperGuy13/crawler-signal-reference.text-retrievaln<1K0 likes35 downloads4mo agoHugging Face24factored /fr-crawler-private-mlm Dataset Card for "fr-crawler-private-mlm" More Information needed text1K<n<10K0 likes34 downloads3y agoHugging Face25factored /fr_crawler_mlm Dataset Card for "fr_crawler" More Information needed text100K<n<1M0 likes31 downloads3y agoHugging Face26factored /fr_crawler_class Dataset Card for "fr_crawler2" More Information needed text1M<n<10M0 likes31 downloads3y agoHugging Face27pixeloffice /ai-crawlers-ip-and-user-agent-registry Known AI Crawlers, User-Agents & Verified IP Ranges Registry (2026) Official aggregated registry of all major AI Scrapers, Search Bots, and LLM Crawlers operating on the web. It includes verified User-Agent strings, policy block targets (for robots.txt), and official CIDR IP ranges for firewall configuration. Published by Pixel Office EU. Purpose & Utility As AI web indexing scales exponentially, website administrators and system engineers face the challenge of… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/ai-crawlers-ip-and-user-agent-registry.texttext-classificationn<1K0 likes29 downloads1mo agoHugging Face28phucnn /zalo-crawler-v17-explanation Dataset Card for "zalo-crawler-v17-explanation" More Information needed text100K<n<1M0 likes27 downloads3y agoHugging Face29logiover /sitemap-to-url-crawler-sample-data Sitemap to URL Crawler nstantly extract all public URLs from any website's sitemap.xml recursively. Handles nested sitemap indexes automatically. The fastest & cheapest way to build URL lists for RAG pipelines, LLM training, and SEO audits. Zero-config & blazing fast. What the actor scrapes Sitemap to URL Crawler — RAG & AI Data Feeder Extract every public URL from any website's sitemap.xml — recursively, instantly, and at scale. Handles nested sitemap indexes… See the full description on the dataset page: https://huggingface.co/datasets/logiover/sitemap-to-url-crawler-sample-data.textn<1K0 likes24 downloads4mo agoHugging Face30IaraMed /crawler_v0text10K<n<100K0 likes17 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.