CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ARKseal /YFCC14M_subset_webdatasetimage1M<n<10M0 likes393 downloads5y agoHugging Face02BEE-spoke-data /open-web-math-minhash Dataset Card for "open-web-math-minhash" An attempt at a "high quality sample" of open-web-math/open-web-math by aggressively applying minhash from text-dedup. The result is 1.82M rows down from the original 6M: DatasetDict({ train: Dataset({ features: ['url', 'text', 'date', 'metadata'], num_rows: 1820241 }) }) Usage Unless you need the metadata, load the text-only config which is only 1.4 GB/5 shards: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/open-web-math-minhash.texttext-generation1M<n<10M0 likes231 downloads9mo agoHugging Face03Lots-of-LoRAs /task1728_web_nlg_data_to_text Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.texttext-generation1K<n<10K0 likes174 downloads2y agoHugging Face04shichen1231 /GBC1M_webdataimage1M<n<10M0 likes126 downloads2y agoHugging Face05open-athena /glm52-datagen-r11-19-knowledge-web-search-mcqa-tracestext1K<n<10K0 likes95 downloads2mo agoHugging Face06Anish13 /web-agent-graph-dataset Web Agent Grouped Graph Dataset This dataset contains web navigation tasks in grouped graph format with full history and candidate actions for training reward models. Data Format Each line in graph_dataset.jsonl represents a single step with all candidate actions grouped together: { "task_id": "...", "goal": "Find product X and add to cart", "domain": "shopping", "step_index": 3, "history": [ {"state_id": "S0", "screenshot": "...", "url": "...", "obs":… See the full description on the dataset page: https://huggingface.co/datasets/Anish13/web-agent-graph-dataset.reinforcement-learning1K<n<10K0 likes77 downloads5mo agoHugging Face07jed351 /Cantonese-Web-Data Dataset Summary Cantonese has been a low-resource language in NLP. This dataset is a major step towards changing that. To our knowledge, this is the first large-scale, properly curated, and deduplicated web dataset built specifically for Cantonese. It was created by filtering years of Common Crawl data and a Cantonese language detector, followed by a deduplication process using MinHash. The result is a high-quality collection of ~250K unique documents containing ~150 million words.… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Cantonese-Web-Data.text100K<n<1M5 likes76 downloads1y agoHugging Face08MC7ever /web-scraper-datasettext10K<n<100K0 likes58 downloads9h agoHugging Face09vyykaaa /dataset-web-attack-newstext10K<n<100K1 likes49 downloads9mo agoHugging Face10AmanPriyanshu /tool-reasoning-sft-TOOLS-toolmind-web-qa-sft-tool-use-data-cleaned-rectified-5.2k ToolMind-Web-QA — Hermes Reasoning Format Filtered and restructured version of Nanbeige/ToolMind-Web-QA. Filters applied: valid role transitions only · known tools only · non-empty user + answer required Size: 5,274 examples (from 5,624 original trajectories, 350 dropped) Source The original dataset contains 5,624 complex multi-hop QA trajectories grounded in Wikipedia entity-relation graphs. Each trajectory has an average of ~138 turns with multiple tool calls across… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolmind-web-qa-sft-tool-use-data-cleaned-rectified-5.2k.texttext-generation1K<n<10K0 likes37 downloads7mo agoHugging Face11Anish13 /web_prm_processed_dataimage10K<n<100K0 likes36 downloads5mo agoHugging Face12miansuleman /web.facebook.com-ru4Ak2ui-scraped-data-Final-Evaltextn<1K0 likes30 downloads1y agoHugging Face13eagle0504 /youthless-homeless-shelter-web-scrape-datasettextn<1K0 likes27 downloads3y agoHugging Face14eagle0504 /youthless-homeless-shelter-web-scrape-dataset-largetextn<1K0 likes27 downloads3y agoHugging Face15brando /small-open-web-math-dataset-v2# Small Open Web Math Dataset v2 A 10k-sample shuffled subset of OpenWebMath, ensuring randomized selection of high-quality mathematical text. text10K<n<100K1 likes25 downloads2y agoHugging Face16brando /small-open-web-math-dataset# Small Open Web Math Dataset A 10k-sample subset of OpenWebMath, focused on high-quality mathematical text. text10K<n<100K1 likes23 downloads2y agoHugging Face17eagle0504 /larkin-web-scrape-dataset-qa-formattedtextn<1K1 likes21 downloads3y agoHugging Face18electricsheepafrica /africa-web-attack-dataset Web Application Attacks (Africa) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-web-attack-dataset.tabulartabular-classification10K<n<100K0 likes21 downloads1mo agoHugging Face19vyykaaa /dataset-web-attacktext100K<n<1M0 likes19 downloads9mo agoHugging Face20Cartinoe5930 /web_text_synthetic_dataset_50ktext10K<n<100K5 likes18 downloads2y agoHugging Face21jwaters8978 /web_scraper_datasetimage10K<n<100K0 likes18 downloads2y agoHugging Face22omarmohamed /Fine_web_lmsys_datasettext1M<n<10M0 likes18 downloads1y agoHugging Face23alkaren38gmailcom /sap-web-datatext100K<n<1M0 likes18 downloads5mo agoHugging Face24eagle0504 /ysa-web-scrape-dataset-qa-formatted-small-versiontextn<1K1 likes17 downloads3y agoHugging Face25daparasyte /webdatatext1K<n<10K0 likes17 downloads2y agoHugging Face26supergoose /flan_combined_task1728_web_nlg_data_to_texttext10K<n<100K0 likes17 downloads2y agoHugging Face27eagle0504 /larkin-web-scrape-dataset-qa-formatted-small-versiontextn<1K0 likes16 downloads3y agoHugging Face28Markie77 /New-Testament-World-English-Web-Dataset-V1 New-Testament-World-English-Web-Dataset-V1 Made with ❤️ using 🦥 Unsloth Studio New Testament World English Web Dataset V1 was generated with Unsloth Recipe Studio. It contains 700 generated records. 🚀 Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("Markie77/New-Testament-World-English-Web-Dataset-V1", "data", split="train") df = dataset.to_pandas() 📊 Dataset Summary 📈 Records: 700 📋 Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/Markie77/New-Testament-World-English-Web-Dataset-V1.textn<1K0 likes15 downloads1mo agoHugging Face29eagle0504 /youthless-homeless-shelter-web-scrape-dataset-qa-formattedtextn<1K1 likes14 downloads3y agoHugging Face30GruberJakub /ai_web_scraping_datasettext10K<n<100K0 likes14 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.