CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AdaMLLab /WebTerminal Terminal/CLI Web Text A filtered extract of terminal and command-line content from two large web-text corpora, designed for upsampling agentic-adjacent data during pretraining. Subsets Subset Rows Tokens Size Quality clean (default) 2.33M 4.6B 11 GB ~98% terminal content unfiltered 61.3M 359B 962 GB ~15% terminal content from datasets import load_dataset # Load the clean subset (default) ds = load_dataset("AdaMLLab/WebTerminal") # Load the unfiltered… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/WebTerminal.tabulartext-generation10M<n<100M4 likes805 downloads7mo agoHugging Face02Brainquiver /general-web-it-202608 General · Web · Italian · 2026-08 Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 21,065,052 documents and 66,158,573,443 characters of Italian prose. Contents Config Documents Characters Upstream fineweb2-hq-ita_Latn 21,065,052 66,158,573,443 epfml/FineWeb2-HQ, ita_Latn The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.tabulartext-generation10M<n<100M0 likes561 downloads27d agoHugging Face03Brainquiver /general-web-fr-202608 General · Web · French · 2026-08 French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 31,999,309 documents and 118,346,333,763 characters of French prose. Contents Config Documents Characters Upstream fineweb2-hq-fra_Latn 31,999,309 118,346,333,763 epfml/FineWeb2-HQ, fra_Latn The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.tabulartext-generation10M<n<100M0 likes510 downloads27d agoHugging Face04Web3Survivor /Survivor 📚 FinePDFs-Edu 350B+ of highly educational tokens from PDFs 📄 What is it? 📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages. FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset. We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/Web3Survivor/Survivor.tabulartext-generation10M<n<100M2 likes382 downloads10mo agoHugging Face05placeholderlabs /exp-pool-olmo-web-dolma2-tokenized Locus EXP OLMo Web - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes130 downloads1mo agoHugging Face06shaikat005 /medium-web-pentesting Medium Web Pentesting Articles Dataset Description A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body. This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.tabulartext-classificationn<1K0 likes109 downloads5mo agoHugging Face07muzaffercky /kurdish-web Dataset Card for Kurdish Web Corpus (Deduplicated) Dataset Summary A multilingual corpus of web text in three Kurdish/Zaza varieties — Kurmanji Kurdish (kmr_Latn), Sorani Kurdish (ckb_Arab), and Zazaki (diq_Latn) — scraped from Kurdish-language websites, language-identified with GlotLID, and deduplicated (exact + MinHash near-duplicate removal). See the "Dataset Creation" section below for the full pipeline. Rows, by language config: config language script… See the full description on the dataset page: https://huggingface.co/datasets/muzaffercky/kurdish-web.tabulartext-generation100K<n<1M1 likes90 downloads3mo agoHugging Face08ZYao720 /WEBPRMBENCH WebPRMBench The first comprehensive evaluation benchmark for Web Process Reward Models Published at ICLR 2026 Paper | Code | Website | Collection | Demo Overview WebPRMBench is the first comprehensive evaluation benchmark dedicated to Web Process Reward Models (WebPRMs). It evaluates how well a reward model can judge the quality of web agent actions during long-horizon web navigation. Each instance presents a web state (page context, trajectory history, user… See the full description on the dataset page: https://huggingface.co/datasets/ZYao720/WEBPRMBENCH.tabulartext-generation1K<n<10K2 likes65 downloads6mo agoHugging Face09ray0rf1re /unclean-web 🕸️ Unclean Web Raw, unfiltered web data scraped across a wide variety of sites and packaged for language model pre-training, fine-tuning, and research. 📊 Dataset Statistics Metric Value Total Pages 18,896 Total Token Estimate 24.38M Unique Sources 54 Schema Version 3.0 Last Updated 2026-06-07 00:36 UTC 🗂️ Available Splits / Subsets Split Description Format full (per batch) Complete raw scrape, all columns… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/unclean-web.tabulartext-generation10K<n<100K0 likes62 downloads4mo agoHugging Face10xlelords /webui 🌐 WebUI A high-quality dataset for training AI models to understand and generate modern websites. The dataset contains structured webpage samples including HTML, screenshots, UI element annotations, semantic labels, color palettes, fonts, layout metadata, and accessibility information. Each sample represents a complete webpage that can be used for web generation, UI understanding, or multimodal training. :contentReference[oaicite:0]{index=0} ✨ Features 📄… See the full description on the dataset page: https://huggingface.co/datasets/xlelords/webui.tabulartext-generation1K<n<10K0 likes48 downloads2mo agoHugging Face11Ba2han /mogan-turkish-web-long moganai/mogan-turkish-web long filtered Turkish texts Source: moganai/mogan-turkish-web (config: default, revision: d773a0efd1b7daf72c7909c83dbd385e7e3564d7). Rows contain 4,000–16,000 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected rows: 2,805,010. Generated by process_hf_dataset.py. See summary.json for counts and thresholds. tabulartext-generation1M<n<10M0 likes43 downloads6d agoHugging Face12nampdn-ai /tiny-webtextgated Tiny WebText The Tiny WebText dataset is designed to help models learn about perception on web text while neutralizing the bias of the source text using critical thinking methods. By providing a rich and diverse set of texts, I aim to improve the ability of models to understand and analyze information in a more objective and unbiased manner. This dataset can be used to train and evaluate natural language processing and machine learning models, with the goal of improving their… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-webtext.tabulartext-generation1M<n<10M37 likes30 downloads3y agoHugging Face13axmeeabdhullo /axya-tech-websearch Dhivehi Combined Dataset Overview This dataset combines three 30K Dhivehi language corpora from the Leipzig Corpora Collection into a single unified CSV file containing 90,000 sentences. The dataset provides a comprehensive resource for Dhivehi language processing, combining data from Wikipedia, news sources, and web crawls. Dataset Composition The dataset consists of three distinct sources: Wikipedia (2021): 30,000 sentences from Dhivehi Wikipedia News… See the full description on the dataset page: https://huggingface.co/datasets/axmeeabdhullo/axya-tech-websearch.tabulartext-generation10K<n<100K0 likes17 downloads1y agoHugging Face14soumikmahato /diverse-websearch-3.5k Diverse WebSearch 3.5k Diverse WebSearch 3.5K is a small web research dataset containing 3,500 diverse web pages collected from search results. Each row contains a source URL, extracted markdown content, a concise summary, and image URLs found on the page. Dataset Details This dataset is intended for learning and experimentation with: webpage summarization retrieval-augmented generation search result understanding document cleaning synthetic QA generation dataset… See the full description on the dataset page: https://huggingface.co/datasets/soumikmahato/diverse-websearch-3.5k.tabularsummarization1K<n<10K1 likes15 downloads5mo agoHugging Face15SPAISS6F1 /spai-ss6-corpus-medical-health-web SPAI SS6 Thai Medical Health Web Corpus Thai public medical and health web articles collected by the local scraping pipeline. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: default Rows in canonical config: 3,660 Parquet size in canonical config: 0.01 GB Source license: other… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-health-web.tabulartext-generationn<1K0 likes11 downloads4mo agoHugging Face16SPAISS6F1 /spai-ss6-corpus-wangchanlion-web SPAI SS6 WangchanLION Web Corpus Index Index repo for the WangchanLION-Web corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: wangchanlion_web Rows in canonical config: 557,502 Parquet size in canonical config: 1.97 GB Source license: odc-by… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-wangchanlion-web.tabulartext-generationn<1K0 likes11 downloads4mo agoHugging Face17web3w /role-play-bench Role-play Benchmark A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios. Dataset Summary Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?". Instead… See the full description on the dataset page: https://huggingface.co/datasets/web3w/role-play-bench.tabulartext-generation1K<n<10K0 likes10 downloads8mo agoHugging Face18OptiTransferData /swiss-web-premium-chgated *.ch Swiss Web Premium (A+) Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- Full provenance -- PII-redacted -- RAG-ready -- SFT-formatted A production-grade Swiss web corpus from the .ch TLD namespace. 110,491 documents independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Built for LLM training, RAG pipelines, SFT fine-tuning, and multilingual NLP. OptiTransferData Portfolio Premium sovereign web corpora for… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch.documenttext-generation10K<n<100K0 likes10 downloads6mo agoHugging Face19weblab-GENIAC /OpenBookQA-Japanese-maskedgated OpenBookQA-Japanese-masked 与えられた問題に対して4つの選択肢から答えを選択するデータセット allenai/openbookqaをcyberagent/calm3-22b-chatで翻訳 5,957件 train split: 4,956件(4,957件の内1件削除) validation split: 500件 test split: 499件(500件の内1件削除) ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施 Format データセットの構成は以下 { "idx": ID, "id": 元ID, "question_stem_en": 英語の質問文, "choices_en": { "text": 選択肢の文章, "label": 選択肢の記号, }… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/OpenBookQA-Japanese-masked.tabulartext-generation1K<n<10K0 likes6 downloads2y agoHugging Face20OptiTransferData /swiss-web-premium-ch-fullgated *.ch Swiss Web Premium (A+) -- Full Dataset Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- 22 production files -- 554.1 MB The complete production release of OptiTransfer's Swiss web corpus. 110,491 documents from the .ch TLD namespace, independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Delivered in Parquet, JSONL, language splits, and pre-built RAG chunks. This is the full commercial dataset. For evaluation, see the free… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch-full.tabulartext-generation100K<n<1M0 likes5 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.