CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ARKseal /YFCC14M_subset_webdatasetimage1M<n<10M0 likes393 downloads5y agoHugging Face02BEE-spoke-data /open-web-math-minhash Dataset Card for "open-web-math-minhash" An attempt at a "high quality sample" of open-web-math/open-web-math by aggressively applying minhash from text-dedup. The result is 1.82M rows down from the original 6M: DatasetDict({ train: Dataset({ features: ['url', 'text', 'date', 'metadata'], num_rows: 1820241 }) }) Usage Unless you need the metadata, load the text-only config which is only 1.4 GB/5 shards: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/open-web-math-minhash.texttext-generation1M<n<10M0 likes231 downloads9mo agoHugging Face03Lots-of-LoRAs /task1728_web_nlg_data_to_text Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.texttext-generation1K<n<10K0 likes174 downloads2y agoHugging Face04shichen1231 /GBC1M_webdataimage1M<n<10M0 likes126 downloads2y agoHugging Face05open-athena /glm52-datagen-r11-19-knowledge-web-search-mcqa-tracestext1K<n<10K0 likes95 downloads2mo agoHugging Face06jed351 /Cantonese-Web-Data Dataset Summary Cantonese has been a low-resource language in NLP. This dataset is a major step towards changing that. To our knowledge, this is the first large-scale, properly curated, and deduplicated web dataset built specifically for Cantonese. It was created by filtering years of Common Crawl data and a Cantonese language detector, followed by a deduplication process using MinHash. The result is a high-quality collection of ~250K unique documents containing ~150 million words.… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Cantonese-Web-Data.text100K<n<1M5 likes76 downloads1y agoHugging Face07MC7ever /web-scraper-datasettext10K<n<100K0 likes58 downloads11h agoHugging Face08vyykaaa /dataset-web-attack-newstext10K<n<100K1 likes49 downloads9mo agoHugging Face09AmanPriyanshu /tool-reasoning-sft-TOOLS-toolmind-web-qa-sft-tool-use-data-cleaned-rectified-5.2k ToolMind-Web-QA — Hermes Reasoning Format Filtered and restructured version of Nanbeige/ToolMind-Web-QA. Filters applied: valid role transitions only · known tools only · non-empty user + answer required Size: 5,274 examples (from 5,624 original trajectories, 350 dropped) Source The original dataset contains 5,624 complex multi-hop QA trajectories grounded in Wikipedia entity-relation graphs. Each trajectory has an average of ~138 turns with multiple tool calls across… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolmind-web-qa-sft-tool-use-data-cleaned-rectified-5.2k.texttext-generation1K<n<10K0 likes37 downloads7mo agoHugging Face10Anish13 /web_prm_processed_dataimage10K<n<100K0 likes36 downloads5mo agoHugging Face11miansuleman /web.facebook.com-ru4Ak2ui-scraped-data-Final-Evaltextn<1K0 likes30 downloads1y agoHugging Face12eagle0504 /youthless-homeless-shelter-web-scrape-datasettextn<1K0 likes27 downloads3y agoHugging Face13eagle0504 /youthless-homeless-shelter-web-scrape-dataset-largetextn<1K0 likes27 downloads3y agoHugging Face14brando /small-open-web-math-dataset-v2# Small Open Web Math Dataset v2 A 10k-sample shuffled subset of OpenWebMath, ensuring randomized selection of high-quality mathematical text. text10K<n<100K1 likes25 downloads2y agoHugging Face15brando /small-open-web-math-dataset# Small Open Web Math Dataset A 10k-sample subset of OpenWebMath, focused on high-quality mathematical text. text10K<n<100K1 likes23 downloads2y agoHugging Face16eagle0504 /larkin-web-scrape-dataset-qa-formattedtextn<1K1 likes21 downloads3y agoHugging Face17electricsheepafrica /africa-web-attack-dataset Web Application Attacks (Africa) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-web-attack-dataset.tabulartabular-classification10K<n<100K0 likes21 downloads1mo agoHugging Face18vyykaaa /dataset-web-attacktext100K<n<1M0 likes19 downloads9mo agoHugging Face19Cartinoe5930 /web_text_synthetic_dataset_50ktext10K<n<100K5 likes18 downloads2y agoHugging Face20jwaters8978 /web_scraper_datasetimage10K<n<100K0 likes18 downloads2y agoHugging Face21omarmohamed /Fine_web_lmsys_datasettext1M<n<10M0 likes18 downloads1y agoHugging Face22alkaren38gmailcom /sap-web-datatext100K<n<1M0 likes18 downloads5mo agoHugging Face23eagle0504 /ysa-web-scrape-dataset-qa-formatted-small-versiontextn<1K1 likes17 downloads3y agoHugging Face24daparasyte /webdatatext1K<n<10K0 likes17 downloads2y agoHugging Face25supergoose /flan_combined_task1728_web_nlg_data_to_texttext10K<n<100K0 likes17 downloads2y agoHugging Face26eagle0504 /larkin-web-scrape-dataset-qa-formatted-small-versiontextn<1K0 likes16 downloads3y agoHugging Face27Markie77 /New-Testament-World-English-Web-Dataset-V1 New-Testament-World-English-Web-Dataset-V1 Made with ❤️ using 🦥 Unsloth Studio New Testament World English Web Dataset V1 was generated with Unsloth Recipe Studio. It contains 700 generated records. 🚀 Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("Markie77/New-Testament-World-English-Web-Dataset-V1", "data", split="train") df = dataset.to_pandas() 📊 Dataset Summary 📈 Records: 700 📋 Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/Markie77/New-Testament-World-English-Web-Dataset-V1.textn<1K0 likes15 downloads1mo agoHugging Face28eagle0504 /youthless-homeless-shelter-web-scrape-dataset-qa-formattedtextn<1K1 likes14 downloads3y agoHugging Face29GruberJakub /ai_web_scraping_datasettext10K<n<100K0 likes14 downloads6mo agoHugging Face30electricsheepafrica /africa-dark-web-data-trading Dark Web African Data Trading | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: parquet, optimized-parquet - Sector: governance_security - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-dark-web-data-trading.tabulartabular-classification10K<n<100K0 likes13 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.