CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01YWZBrandon /webshop-data4 likes94k downloads1y agoHugging Face02EssentialAI /essential-web-v1.0 🌐 Essential-Web: Complete 24-Trillion Token Dataset 🏆 Website | 🖥️ Code | 📖 Paper | ☁️ AWS 📋 Dataset Description Essential-Web is a 24-trillion-token web dataset with document-level metadata designed for flexible dataset curation. The dataset provides metadata including subject matter classification, web page type, content complexity, and document quality scores for each of the 23.6 billion documents. Researchers can filter and curate specialized datasets… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/essential-web-v1.0.10B<n<100B247 likes69k downloads1y agoHugging Face03lucas-ventura /WebVidvideo1K<n<10K1 likes60k downloads5mo agoHugging Face04WebOrganizer /Corpus-200B WebOrganizer/Corpus-200B [Paper] [Website] [GitHub] This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned with (1) RefinedWeb filters and (2) BFF deduplication. We provide the resulting 200B token corpus annotated with two quality scores, WebOrganizer domains, and k-means scores. Download the dataset by cloning the repository with Git LFS instead of HuggingFace's load_dataset(). The dataset has the following folder structure:… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/Corpus-200B.text100M<n<1B12 likes39k downloads3mo agoHugging Face05McGill-NLP /weblinx-browsergym WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the video tag. This dataset was specifically created to allow WebLINX to be used inside the BrowserGym and Agentlab ecosystem. Please see the browsergym repository for more information. [!NOTE] The version associated with this library is WebLINX… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/weblinx-browsergym.image-to-text4 likes34k downloads2y agoHugging Face06open-web-math /open-web-math Keiran Paster*, Marco Dos Santos*, Zhangir Azerbayev, Jimmy Ba GitHub | ArXiv | PDF OpenWebMath is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of 6.3 million documents containing a total of 14.7B tokens. OpenWebMath is intended for use in pretraining and finetuninglarge language models. You can download the dataset using Hugging Face: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/open-web-math/open-web-math.text1M<n<10M361 likes33k downloads3y agoHugging Face07McGill-NLP /WebLINX-full WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead: WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models 💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.text10K<n<100K8 likes28k downloads1y agoHugging Face08prasadonly /webtoepub-library8 likes27k downloads1m agoHugging Face09HuggingFaceM4 /WebSight Dataset Card for WebSight Dataset Description WebSight is a large synthetic dataset containing HTML/CSS codes representing synthetically generated English websites, each accompanied by a corresponding screenshot. This dataset serves as a valuable resource for tasks such as generating UI codes from a screenshot. It comes in two versions: v0.1: Websites are coded with HTML + CSS. They do not include real images. v0.2: Websites are coded with HTML + Tailwind CSS. They do… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/WebSight.image1M<n<10M400 likes18k downloads2y agoHugging Face10birder-project /TreeOfLife-10M-WEBP Dataset Card for TreeOfLife-10M-WEBP Dataset Description This is an optimized version of the TreeOfLife-10M dataset, containing over 10 million images covering 454 thousand taxa in the tree of life. This version has been processed to improve usability and reduce storage requirements while maintaining full compatibility with the original dataset structure. Dataset Summary This version modifies the original dataset as follows: Corrupted files were… See the full description on the dataset page: https://huggingface.co/datasets/birder-project/TreeOfLife-10M-WEBP.image-classification10M<n<100M1 likes17k downloads2mo agoHugging Face11stanfordnlp /web_questions Dataset Card for "web_questions" Dataset Summary This dataset consists of 6,642 question/answer pairs. The questions are supposed to be answerable by Freebase, a large knowledge graph. The questions are mostly centered around a single named entity. The questions are popular ones asked on the web (at least in 2013). Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/web_questions.textquestion-answering1K<n<10K42 likes13k downloads3y agoHugging Face12TempoFunk /webvid-10Mtexttext-to-video10M<n<100M98 likes11k downloads3y agoHugging Face13BIOMEDICA /biomedica_webdataset_24Mgated Dataset Card for Dataset Name Arxiv: Arxiv &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Website: Biomedica &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Training instructions: OpenCLIP &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Tutorial: Google Colab BIOMEDICA Dataset is a large-scale, deep-learning-ready biomedical dataset containing over 24M imagecaption pairs and 30M image-references from 6M unique open-source articles. Each… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/biomedica_webdataset_24M.n>1T40 likes9.2k downloads20d agoHugging Face14Research-EAI /essential-web-1t-sample-fdc-partitioned 🌐 Essential-Web: FDC Level-2 Partitioned Dataset 📋 Dataset Description This dataset contains a 1 trillion token sample from Essential-Web, partitioned by Free Decimal Correspondence (FDC) level-2 categories. Essential-Web is a 24-trillion-token web dataset with extensive document-level metadata designed to enable rapid dataset curation through SQL-like filtering. 🔍 Free Decimal Correspondence (FDC) The FDC taxonomy is an open classification system… See the full description on the dataset page: https://huggingface.co/datasets/Research-EAI/essential-web-1t-sample-fdc-partitioned.text100M<n<1B5 likes8.9k downloads1y agoHugging Face15WorldEngineAI /WEB-Datasetgated WorldEngine Bimanual Dataset for Post-training A large-scale, language-annotated real-robot bimanual manipulation dataset for post-training robotics foundation models. It spans 90 everyday manipulation tasks collected with a bimanual YAM follower arm teleoperated by a GELLO leader, recording joint state, action, and three synchronized camera streams at 60 Hz. Shared lineage, different story. This dataset shares its hardware, teleoperation setup, and recording pipeline with the… See the full description on the dataset page: https://huggingface.co/datasets/WorldEngineAI/WEB-Dataset.video100K<n<1M1 likes8.2k downloads2mo agoHugging Face16deepghs /safebooru-webp-4Mpixelgated Safebooru 4M Re-encoded Dataset This is the re-encoded dataset of deepghs/safebooru_full. And all the resized images are maintained here. There are 5756655 images in total. The maximum ID of these images is 5974383. Last updated at 2025-08-06 08:31:53 JST. How to Painlessly Use This Use cheesechaser to quickly get images from this repository. Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/safebooru-webp-4Mpixel.image-classification1M<n<10M2 likes7.6k downloads1y agoHugging Face17farama-minari /webagents0 likes7k downloads9mo agoHugging Face18PaDaS-Lab /webfaq-retrievalWebFAQ Retrieval Dataset Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Retrieval Dataset is a carefully filtered and curated subset of the broader WebFAQ Q&A Dataset.It is purpose-built for Information Retrieval (IR) tasks, such as training and evaluating dense or sparse retrieval models in multiple languages. Each of the… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval.texttext-retrieval10M<n<100M10 likes6.6k downloads1y agoHugging Face19ArneH /swiss-caselaw-web-ui Swiss Case Law Open Dataset 962,724 published decisions from Swiss federal, cantonal, and regulatory bodies. Full text, structured metadata, and daily updates. The March 20, 2026 snapshot contains German, French, and Italian decisions; the export schema also reserves rm for Romansh. What this is A structured, searchable archive of Swiss court decisions — from the Federal Supreme Court (BGer) down to cantonal courts in all 26 cantons. Every decision includes the full… See the full description on the dataset page: https://huggingface.co/datasets/ArneH/swiss-caselaw-web-ui.0 likes6.4k downloads6mo agoHugging Face20laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes6.2k downloads4y agoHugging Face21yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M52 likes6.2k downloads6mo agoHugging Face22TIGER-Lab /WebInstructSub 🦣 MAmmoTH2: Scaling Instructions from the Web Project Page: https://tiger-ai-lab.github.io/MAmmoTH2/ Paper: https://arxiv.org/pdf/2405.03548 Code: https://github.com/TIGER-AI-Lab/MAmmoTH2 WebInstruct (Subset) This repo contains the partial dataset used in "MAmmoTH2: Scaling Instructions from the Web". This partial data is coming mostly from the forums like stackexchange. This subset contains very high-quality data to boost LLM performance through instruction tuning.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/WebInstructSub.textquestion-answering1M<n<10M164 likes6.2k downloads2y agoHugging Face23xcodemind /webcode2m_purifiedWebCode2M: A Real-World Dataset for Code Generation from Webpage Designs Features: image: the screenshot of the webpage. bbox: the layout information, i.e., the bounding boxes (Bbox) of all the elements in the webpage, which contains the size, position, and hierarchy information. text: the webpage code text including HTML/CSS code. scale: the scale of the screenshot, in the format [width, height]. lang: the main language of the text content displayed on the rendered page (excluding HTML/CSS… See the full description on the dataset page: https://huggingface.co/datasets/xcodemind/webcode2m_purified.imageimage-to-text1M<n<10M6 likes5.1k downloads2y agoHugging Face24OpenTransformer /web-crawl-2026 Web Crawl 2026 A large-scale web crawl dataset for language model pretraining, collected by the OpenTransformer project. Dataset Description This dataset contains text extracted from web pages crawled directly from the internet using custom high-throughput crawlers. All data is freshly scraped. Data Format Each record is a JSON line (gzipped) with fields: text: extracted text content (200-200,000 chars) url: source URL domain: source domain timestamp: crawl… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-2026.text-generation10B<n<100B1 likes5k downloads5mo agoHugging Face25webshart /conceptual-captions-12m-webdataset-metadata Conceptual Captions 12M — Webshart metadata indices Per-shard webshart metadata indices for laion/conceptual-captions-12m-webdataset: 1,100 JSON files under data/, one per source tar shard, mirroring the source's shard layout. Each index records every tar member's byte offset and length (enabling ranged reads without downloading whole shards), image geometry (width/height for aspect bucketing), and — as of August 2026 — embedded captions for all 10,994,853 samples, coalesced… See the full description on the dataset page: https://huggingface.co/datasets/webshart/conceptual-captions-12m-webdataset-metadata.1 likes4.8k downloads1mo agoHugging Face26lee101 /webfiddle-internet-raw-cache-datasetA dataset of different files that robots tried to crawl through webfiddle.net Mostly html files but other files too pdfs, images, binary- i have no idea what is in here at this stage - but gives an interesting idea of what crawlers like to visit and could be the basis of interesting SEO or coding LLM reasearch. Collected as part of my work on web simulators. https://webfiddle.net JS/CSS editor for the web, https://websim.netwrck.com Coding Editor for the web. https://x.com/leeleepenkman Its… See the full description on the dataset page: https://huggingface.co/datasets/lee101/webfiddle-internet-raw-cache-dataset.2 likes4.7k downloads0m agoHugging Face27microsoft /webgym_tasks WebGym Tasks Dataset Dataset Description This dataset contains web navigation tasks for training and evaluating autonomous web agents. Each task consists of a natural language instruction that describes an action to be performed on a specific website, along with evaluation criteria and metadata. Dataset Summary Total Training Tasks: 292,092 Total Test Tasks: 1,167 Domains: Multiple domains including Lifestyle & Leisure, Sports & Fitness, and more Source… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/webgym_tasks.textreinforcement-learning100K<n<1M20 likes4.7k downloads7mo agoHugging Face28mteb /webis-touche2020-v3 Touche2020Retrieval.v3 An MTEB dataset Massive Text Embedding Benchmark Touché Task 1: Argument Retrieval for Controversial Questions Task category t2t Domains Academic Reference https://github.com/castorini/touche-error-analysis How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["Touche2020Retrieval.v3"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/webis-touche2020-v3.texttext-retrieval100K<n<1M0 likes4.4k downloads1y agoHugging Face29rmanluo /RoG-webqsp Dataset Card for "RoG-webqsp" More Information needed text1K<n<10K29 likes4.3k downloads3y agoHugging Face30OctoThinker /MegaMath-Web-Pro-Max OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Curation of MegaMath-Web-Pro-Max Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year; Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data; Step 3: Training a fasttext carefully with proper preprocessing; Step 4: Filtering documents with a threshold (i.e., 0.4); Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.tabular10M<n<100M41 likes4.3k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.