CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hudsongouge /AoPS-Scrape AoPS-Scrape Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints. Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session. Splits Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split: Split Rows Notes deduplicated 29,964 One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.tabularquestion-answering10K<n<100K2 likes4.2k downloads2mo agoHugging Face02Gugu8 /Scraped-Datagated Scraped-Data — Building to 1 TB A continuously-growing, cleaned text corpus scraped, parsed, cleaned, losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline. Current repo size: ~26 GB → target: 1 TB (incremental batch uploads). Current files File Size Contents corpus_batch1.clean.zip 4.80 GB ~2.9M Wikipedia docs (stored, from first scrape) corpus_batch2.clean.zip 4.85 GB ~2.5M docs corpus_batch3.clean.zip 4.91 GB ~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.texttext-generation100M<n<1B0 likes1.8k downloads17d agoHugging Face03thonyyy /tatoeba-nusax-scrape-mt-concattext100M<n<1B0 likes585 downloads2y agoHugging Face04Amin1600 /Web_Scraper_Datatext10K<n<100K1 likes563 downloads18h agoHugging Face05mkd-minju /korean_data_scraper_aihub Korean Data Scraper — AI-Hub Local Corpus 상태: 비공개 (private) 저장소입니다. Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다. 스키마 파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다. {"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.texttext-generation1M<n<10M0 likes352 downloads22d agoHugging Face06endomorphosis /legal_scrapers JusticeDAO Legal Scrapers Research collectors that download official legislative sources only (national gazettes and official open-data portals). Intended for building reproducible legal-text corpora, not for production legal research products. Not legal advice. The official gazette of each jurisdiction prevails. License AGPL-3.0 for the collector scripts in this repository. Contents scrapers/ includes country collectors, EU/EUR-Lex, shared helpers… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/legal_scrapers.texttext-retrievaln<1K0 likes326 downloads14d agoHugging Face07logiover /gleif-lei-scraper-sample-data GLEIF LEI Scraper Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run. What the actor scrapes 🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.textn<1K0 likes280 downloads4mo agoHugging Face08Isamu136 /penetration_testing_scraped_dataset Dataset Card for "penetration_testing_scraped_dataset" More Information needed text100K<n<1M14 likes243 downloads3y agoHugging Face09ryang2 /linkedin-job-scrape LinkedIn DS/ML Job Postings Daily snapshots of data science / machine learning job postings scraped from LinkedIn, one parquet file per scrape run. All splits share one schema (2023 → today); historical splits were migrated in 2026-07 — the original 8-column data is preserved at revision v1-schema. The same job_id recurs across splits (a posting stays live for days) — that's the longitudinal signal. For a unique-jobs view, dedup on job_id keeping the row with max scrape_dt.… See the full description on the dataset page: https://huggingface.co/datasets/ryang2/linkedin-job-scrape.text100K<n<1M0 likes230 downloads18d agoHugging Face10DarkforumsHeiter /Darkforums-Scraped-Dump-osint-Datatextn<1K1 likes203 downloads3mo agoHugging Face11barryallen16 /fitcheck-scraped-multiviewimage100K<n<1M0 likes199 downloads8mo agoHugging Face12barryallen16 /fitcheck-scraped-v1image100K<n<1M0 likes181 downloads8mo agoHugging Face13trentmkelly /DrugHub-scrape DrugHub Market Snapshot, September 2026 A complete, text-only capture of the public listing, vendor, and review pages of DrugHub, a Monero-only darknet market operating since 2023. Everything here was visible to any visitor without an account. Doesn't include any images. Collected 16-17 September 2026. Enriched with model-derived labels (typesafe/jev-1.13) on 19 September 2026; see the listing_enrichment table and the Enrichment section below. What's in it… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/DrugHub-scrape.tabular100K<n<1M3 likes158 downloads3d agoHugging Face14scrapegraphai /scrapegraphai-100k ScrapeGraphAI-100k Dataset Summary ScrapeGraphAI-100k is a dataset of 93,695 real-world schema-constrained extraction events collected via opt-in telemetry of the ScrapeGraphAI open-source scraping library in Q2–Q3 2025. The dataset was derived from ~9 million raw PostHog telemetry events, deduplicated and balanced for schema diversity. Each example captures one LLM attempt to extract structured data from real web content under a user-defined JSON schema: the… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraphai-100k.tabulartext-generation10K<n<100K26 likes155 downloads2mo agoHugging Face15reapxdev /finra-brokercheck-scraper FINRA BrokerCheck Scraper · Advisors, Firms & Disclosures Scrape financial advisors, firm affiliations, CRDs, registration scope, and disclosure histories directly from FINRA BrokerCheck API into clean dataset rows. Rows in this dataset 1,430 Fields 22 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/finra-brokercheck-scraper.tabular1K<n<10K0 likes131 downloads2mo agoHugging Face16Nicolas-BZRD /English_French_Webpages_Scraped_Translated English French Webpages Scraped Translated Dataset Summary French/English parallel texts for training translation models. Over 17.1 million sentences in French and English. Dataset created by Chris Callison-Burch, who crawled millions of web pages and then used a set of simple heuristics to transform French URLs onto English URLs, and assumed that these documents are translations of each other. This is the main dataset of Workshop on Statistical Machine Translation (WML)… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Webpages_Scraped_Translated.texttranslation10M<n<100M3 likes109 downloads3y agoHugging Face17neptun-org /neptun.scraper Data in this dataset Docker & NPM Scraped using crawl4ai. The NPM and Docker data was scraped from docs.docker.com and docs.npmjs.com and processed using GPT-4 resulting in docker_documentation.jsonl and npm_documentation.jsonl. The file training-data-v1.jsonl also includes Titanium, dockerNLcommands and docker_ps. GitHub Scraped using firecrawl. The GitHub data was scraped from docs.github.com/en using firecrawl.A few pages might be missing in the… See the full description on the dataset page: https://huggingface.co/datasets/neptun-org/neptun.scraper.textquestion-answering100K<n<1M1 likes108 downloads2y agoHugging Face18reapxdev /greenhouse-jobs-scraper Greenhouse Jobs Scraper Scrape every public job posting from any Greenhouse company job board: title, department, location, remote flag, seniority, advertised salary, full description and apply URL. Rows in this dataset 14,091 Fields 36 Collector runs behind it 61 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/greenhouse-jobs-scraper/ — 92 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/greenhouse-jobs-scraper.tabular10K<n<100K0 likes100 downloads2mo agoHugging Face19alxfgh /Protocol_Scrape Dataset Card for "Protocol_Scrape" More Information needed text1K<n<10K0 likes97 downloads2y agoHugging Face20scrapegraphai /scrapegraph-100k-finetuning ScrapeGraphAI 100k finetuning Dataset Summary A finetuning-ready derivative of ScrapeGraphAI-100k: schema-constrained web extraction examples where a model must produce JSON conforming to a user-defined JSON schema given Markdown-converted page content. Split Rows Targets train 25,244 GPT-5-nano regenerated targets test 2,808 GPT-5-nano regenerated targets human_eval 100 Human labeled extractions (evaluation only) Important: train/test… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraph-100k-finetuning.textfeature-extraction10K<n<100K2 likes94 downloads2mo agoHugging Face21mondk /arena.ai-code-leaderboard-scrapedlink: https://arena.ai/leaderboard/code/ tabularn<1K0 likes91 downloads25d agoHugging Face22logiover /steam-game-reviews-scraper-sample-data Steam Game & Reviews Scraper Scrape Steam game metadata, pricing, genres, Metacritic scores & user reviews using Steam's public API. Supports bulk app IDs, store URLs & keyword search. No proxy needed. What the actor scrapes Steam Game & Reviews Scraper — Steam Store Data & User Reviews to JSON/CSV Scrape game metadata and user reviews from the Steam Store using Steam's public JSON API. This Steam scraper extracts prices, discounts, genres, Metacritic scores… See the full description on the dataset page: https://huggingface.co/datasets/logiover/steam-game-reviews-scraper-sample-data.tabularn<1K0 likes85 downloads4mo agoHugging Face23firecrawl /scrape-content-dataset-v1 Scrape Content Dataset v1 A human-curated benchmark dataset for evaluating web scraping engines on content quality. Overview This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time. Dataset Structure CSV format with columns: id: Sequential identifier url:… See the full description on the dataset page: https://huggingface.co/datasets/firecrawl/scrape-content-dataset-v1.text1K<n<10K0 likes83 downloads11mo agoHugging Face24IrishHotelReviews /scraped-hotel-reviewstabular10K<n<100K0 likes73 downloads3mo agoHugging Face25reapxdev /grants-gov-scraper Grants.gov Scraper · Grant Opportunities, Agencies & Awards Scrape US federal grant opportunities, funding announcements, and agency award notices from Grants.gov by keyword, agency, category, eligibility, and status. Rows in this dataset 1,488 Fields 12 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/grants-gov-scraper.text1K<n<10K0 likes72 downloads2mo agoHugging Face26jkorsvik /nowiki_second_scrape_merged Dataset Card for "nowiki_second_scrape_merged" Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/jkorsvik/nowiki_second_scrape_merged.tabularsummarization100K<n<1M0 likes66 downloads4y agoHugging Face27sayurio /ekpatagolpo-scrape-bangla-literature Ekpatagolpo Bengali Stories Archive Request More ScrapesOrder Private Scrapes Overview This repository contains a large-scale, curated text dataset scraped from ekpatagolpo.com. The primary goal of this archive is to preserve a massive collection of purely human-written Bengali literature and stories (Bangla Golpo), creating a distinct record of human creativity separate from AI-generated text. Purpose and Usage This dataset is published publicly under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/ekpatagolpo-scrape-bangla-literature.texttext-generation10K<n<100K1 likes66 downloads6mo agoHugging Face28logiover /defillama-protocols-scraper-sample-data DefiLlama Protocols Scraper Scrape all 7,000+ DeFi protocols from DefiLlama in one run — TVL, 1h/1d/7d TVL change, market cap, category, chains and links. Filter by chain, category and TVL. Schedule it daily to track the entire DeFi landscape. What the actor scrapes 🦙 DefiLlama Protocols Scraper — Scrape All DeFi Protocols & TVL Data Scrape all 7,000+ DeFi protocols from DefiLlama in a single run and export them to JSON, CSV or Excel. This DefiLlama scraper… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-protocols-scraper-sample-data.tabularn<1K0 likes66 downloads4mo agoHugging Face29reapxdev /shopify-store-products-scraper Shopify Store Scraper Scrape products, prices, discounts, variants and stock from any Shopify store's public product JSON. No login, no API key, no headless browser. Rows in this dataset 11,720 Fields 33 Collector runs behind it 72 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/shopify-store-products-scraper/ — 1,271 entity pages Run the collector yourself https://apify.com/reapx/shopify-store-products-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-store-products-scraper.image10K<n<100K0 likes65 downloads2mo agoHugging Face30reapxdev /boardgamegeek-scraper BoardGameGeek Scraper · Games, Ratings, Designers & Mechanics Scrape board games, release years, player counts, categories, mechanics, designers, artists, and publishers from BoardGameGeek. HTTP only, pay-per-event pricing. Rows in this dataset 437 Fields 21 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/boardgamegeek-scraper.imagen<1K0 likes65 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.