CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01webagentlab /webchain WebChain v2 A large-scale, human-annotated dataset of real-world web interaction trajectories for training and evaluating web agents. [Paper] [Code] [Dataset] WebChain captures how people complete real tasks on live websites. It is designed for agents that must both identify the correct interface element and reason through a sequence of actions. Each trajectory aligns screenshots, web structure, grounded actions, and reasoning signals instead of treating web navigation as… See the full description on the dataset page: https://huggingface.co/datasets/webagentlab/webchain.tabular1K<n<10K1 likes4.3k downloads28d agoHugging Face02webninjasi /pk-map-statstabular1M<n<10M3 likes3.3k downloads24d agoHugging Face03OctoThinker /MegaMath-Web-Pro-Max OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Curation of MegaMath-Web-Pro-Max Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year; Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data; Step 3: Training a fasttext carefully with proper preprocessing; Step 4: Filtering documents with a threshold (i.e., 0.4); Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.tabular10M<n<100M41 likes2.7k downloads1y agoHugging Face04NoeFlandre /osm-polygon-website-tag OSM Polygon Website Dataset OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files. At a glance Polygons 1,726,474 With extracted text 1,192,980 Words of text 407,685,655 Languages 397 Regional sources 386 / 386 Duplicate objects removed 104,927 Candidates rejected 868,905,743 Status In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.tabular1M<n<10M2 likes2.4k downloads3d agoHugging Face05AdaMLLab /WebTerminal Terminal/CLI Web Text A filtered extract of terminal and command-line content from two large web-text corpora, designed for upsampling agentic-adjacent data during pretraining. Subsets Subset Rows Tokens Size Quality clean (default) 2.33M 4.6B 11 GB ~98% terminal content unfiltered 61.3M 359B 962 GB ~15% terminal content from datasets import load_dataset # Load the clean subset (default) ds = load_dataset("AdaMLLab/WebTerminal") # Load the unfiltered… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/WebTerminal.tabulartext-generation10M<n<100M4 likes805 downloads7mo agoHugging Face06HAERAE-HUB /KOREAN-WEBTEXT KOREAN-WEBTEXT KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources: cc100 oscar-corpus/OSCAR-2201 oscar-corpus/OSCAR-2109 oscar-corpus/OSCAR-2301 ontocord/CulturaY Additional credible internet sources collected by out team (We are working to add more sources) The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.tabular1M<n<10M49 likes709 downloads2y agoHugging Face07Brainquiver /general-web-it-202608 General · Web · Italian · 2026-08 Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 21,065,052 documents and 66,158,573,443 characters of Italian prose. Contents Config Documents Characters Upstream fineweb2-hq-ita_Latn 21,065,052 66,158,573,443 epfml/FineWeb2-HQ, ita_Latn The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.tabulartext-generation10M<n<100M0 likes561 downloads26d agoHugging Face08Brainquiver /general-web-fr-202608 General · Web · French · 2026-08 French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 31,999,309 documents and 118,346,333,763 characters of French prose. Contents Config Documents Characters Upstream fineweb2-hq-fra_Latn 31,999,309 118,346,333,763 epfml/FineWeb2-HQ, fra_Latn The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.tabulartext-generation10M<n<100M0 likes510 downloads26d agoHugging Face09placeholderlabs /pretrain-web-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 8,689,580,607 (8.7B) Trainable tokens 8,689,580,607 (8.7B) Documents 281,846 Shards 89 UTF-8 bytes 37,540,769,483 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix-long-context.tabular100K<n<1M0 likes469 downloads13d agoHugging Face10lightonai /webis-touche2020-decontaminated webis-touche2020 (Decontaminated) A decontaminated version of the webis-touche2020 dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed. Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/webis-touche2020-decontaminated.tabulartext-retrieval100K<n<1M0 likes448 downloads6mo agoHugging Face11AmineHA /WebArena-Verified WebArena-Verified Dataset description WebArena-Verified is a curated benchmark dataset of web tasks designed for reproducible evaluation of web agents across multiple realistic websites. Sources GitHub repository: webarena-verified Original WebArena benchmark: webarena.dev Splits full: 812 rows hard: 258 rows Tasks per site Counts below are task counts grouped by category. Tasks with more than one site are grouped under multi-category… See the full description on the dataset page: https://huggingface.co/datasets/AmineHA/WebArena-Verified.tabular1K<n<10K2 likes394 downloads8mo agoHugging Face12ARKseal /YFCC14M_subset_webdatasetimage1M<n<10M0 likes393 downloads5y agoHugging Face13Web3Survivor /Survivor 📚 FinePDFs-Edu 350B+ of highly educational tokens from PDFs 📄 What is it? 📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages. FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset. We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/Web3Survivor/Survivor.tabulartext-generation10M<n<100M2 likes382 downloads10mo agoHugging Face14lehduong /megamath-web-protabular10M<n<100M0 likes335 downloads1y agoHugging Face15guanfengliu /so101_main_bin_2cameras_webThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 53, "total_frames": 9844, "total_tasks": 1, "total_videos": 106, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:53" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/guanfengliu/so101_main_bin_2cameras_web.tabularrobotics10K<n<100K0 likes288 downloads1y agoHugging Face16sanchit-gandhi /cosmopedia_web_textbooks_logprobstabular1M<n<10M0 likes207 downloads2y agoHugging Face17webstep /duplo-pick-and-place-yellow-640x480-3-camerasThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/webstep/duplo-pick-and-place-yellow-640x480-3-cameras.tabularrobotics10K<n<100K0 likes207 downloads10d agoHugging Face18crawlora-net /dead-web-commoncrawl Dead-Web Common Crawl — a longitudinal host-reachability panel (2018–2026) &nbsp;·&nbsp; Hugging Face &nbsp;·&nbsp; Kaggle &nbsp;·&nbsp; License: CC BY 4.0 An open dataset labeling the reachability of 172,959,928 registered domains across 80 monthly Common Crawl archives (CC-MAIN-2018-05 → CC-MAIN-2026-21), built from Common Crawl's robotstxt subset and calibrated against a live re-probe to separate genuinely-dead domains from those merely blocking crawlers or dropped from the… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/dead-web-commoncrawl.tabular100M<n<1B1 likes202 downloads2mo agoHugging Face19WebOrganizer /TopicAnnotations-Llama-3.1-8B WebOrganizer/TopicAnnotations-Llama-3.1-8B [Paper] [Website] [GitHub] This dataset contains 1M web pages annotated with topic labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/TopicClassifier. Dataset Structure Each example contains the following fields: text: The text content of the web page url: The URL of the web page top_choice_index: Index of the most… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-8B.tabular1M<n<10M1 likes197 downloads2y agoHugging Face20VedantPadwal /clean-visual-webarena-classifiedstabular10K<n<100K0 likes194 downloads2y agoHugging Face21axon-rl /webshop_instructionstabular1K<n<10K0 likes190 downloads11mo agoHugging Face22webis /rank-distillmThis dataset contains the training run files from the paper Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-ranking for training queries from MS MARCO passage re-ranked by RankZephyr, a large monoELECTRA model or a large Set-Encoder model. These run files can be used to distill smaller and more efficient models while upholding effectiveness. The files __colbert__msmarco-passage-train-judged.parquet and __bm25__msmarco-passage-train-judged.parquet… See the full description on the dataset page: https://huggingface.co/datasets/webis/rank-distillm.tabular100M<n<1B1 likes176 downloads9mo agoHugging Face23guanfengliu /so101_main_bin_2cameras_web2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 43, "total_frames": 13165, "total_tasks": 1, "total_videos": 86, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:43" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/guanfengliu/so101_main_bin_2cameras_web2.tabularrobotics10K<n<100K0 likes171 downloads1y agoHugging Face24sandhyavs /push_purple_block_webcamtabular10K<n<100K0 likes149 downloads1y agoHugging Face25shaowenchen /webtextqa_zhFrom 2015 to 2016 tabular1M<n<10M1 likes141 downloads3y agoHugging Face26placeholderlabs /pretrain-web-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 73,646,210,143 (73.6B) Trainable tokens 73,646,210,143 (73.6B) Documents 61,059,647 Shards 590 UTF-8 bytes 341,537,872,441 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix.tabular100M<n<1B0 likes141 downloads13d agoHugging Face27philipphager /MSLR-WEB10ktabular10K<n<100K0 likes138 downloads2y agoHugging Face28kdcyberdude /cosmopedia_web_samples_v2_shards_entabular1M<n<10M0 likes134 downloads2y agoHugging Face29placeholderlabs /exp-pool-olmo-web-dolma2-tokenized Locus EXP OLMo Web - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes130 downloads1mo agoHugging Face30ROSCOSMOS /Movie-Poster-WebURL-Dataset-1874-2025 Movie Poster WebURL Dataset 1874–2025 A TMDB-derived metadata index of movie poster WebURLs covering 1874–2025. The dataset contains metadata and external TMDB poster URLs. Poster image binaries are not redistributed in this repository. Data Split: train Rows: 804,304 Format: Parquet Columns: 15 The publication artifact was produced from a larger local TMDB harvest and passed a conservative metadata-based content filtering and post-filter verification process… See the full description on the dataset page: https://huggingface.co/datasets/ROSCOSMOS/Movie-Poster-WebURL-Dataset-1874-2025.image100K<n<1M2 likes127 downloads19d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.