datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AoPS-Scrape
AoPS-Scrape
Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints.
Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session.
Splits
Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split:
Split
Rows
Notes
deduplicated
29,964
One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.Scraped-Data
Scraped-Data — Building to 1 TB
A continuously-growing, cleaned text corpus scraped, parsed, cleaned,
losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline.
Current repo size: ~26 GB → target: 1 TB (incremental batch uploads).
Current files
File
Size
Contents
corpus_batch1.clean.zip
4.80 GB
~2.9M Wikipedia docs (stored, from first scrape)
corpus_batch2.clean.zip
4.85 GB
~2.5M docs
corpus_batch3.clean.zip
4.91 GB
~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.tatoeba-nusax-scrape-mt-concatWeb_Scraper_Datakorean_data_scraper_aihub
Korean Data Scraper — AI-Hub Local Corpus
상태: 비공개 (private) 저장소입니다.
Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다.
스키마
파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다.
{"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.legal_scrapers
JusticeDAO Legal Scrapers
Research collectors that download official legislative sources only
(national gazettes and official open-data portals). Intended for building
reproducible legal-text corpora, not for production legal research products.
Not legal advice. The official gazette of each jurisdiction prevails.
License
AGPL-3.0 for the collector scripts in this repository.
Contents
scrapers/ includes country collectors, EU/EUR-Lex, shared helpers… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/legal_scrapers.gleif-lei-scraper-sample-data
GLEIF LEI Scraper
Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run.
What the actor scrapes
🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.penetration_testing_scraped_dataset
Dataset Card for "penetration_testing_scraped_dataset"
More Information needed
linkedin-job-scrape
LinkedIn DS/ML Job Postings
Daily snapshots of data science / machine learning job postings scraped from LinkedIn, one parquet file per scrape run. All splits share one schema (2023 → today); historical splits were migrated in 2026-07 — the original 8-column data is preserved at revision v1-schema.
The same job_id recurs across splits (a posting stays live for days) — that's the longitudinal signal. For a unique-jobs view, dedup on job_id keeping the row with max scrape_dt.… See the full description on the dataset page: https://huggingface.co/datasets/ryang2/linkedin-job-scrape.Darkforums-Scraped-Dump-osint-Datafitcheck-scraped-multiviewfitcheck-scraped-v1DrugHub-scrape
DrugHub Market Snapshot, September 2026
A complete, text-only capture of the public listing, vendor, and review pages of
DrugHub, a Monero-only darknet market operating since 2023. Everything here was
visible to any visitor without an account. Doesn't include any images.
Collected 16-17 September 2026. Enriched with model-derived labels
(typesafe/jev-1.13) on 19 September 2026; see the listing_enrichment
table and the Enrichment section below.
What's in it… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/DrugHub-scrape.scrapegraphai-100k
ScrapeGraphAI-100k
Dataset Summary
ScrapeGraphAI-100k is a dataset of 93,695 real-world schema-constrained extraction events collected via opt-in telemetry of the ScrapeGraphAI open-source scraping library in Q2–Q3 2025. The dataset was derived from ~9 million raw PostHog telemetry events, deduplicated and balanced for schema diversity. Each example captures one LLM attempt to extract structured data from real web content under a user-defined JSON schema: the… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraphai-100k.finra-brokercheck-scraper
FINRA BrokerCheck Scraper · Advisors, Firms & Disclosures
Scrape financial advisors, firm affiliations, CRDs, registration scope, and disclosure histories directly from FINRA BrokerCheck API into clean dataset rows.
Rows in this dataset
1,430
Fields
22
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/finra-brokercheck-scraper.English_French_Webpages_Scraped_Translated
English French Webpages Scraped Translated
Dataset Summary
French/English parallel texts for training translation models. Over 17.1 million sentences in French and English. Dataset created by Chris Callison-Burch, who crawled millions of web pages and then used a set of simple heuristics to transform French URLs onto English URLs, and assumed that these documents are translations of each other. This is the main dataset of Workshop on Statistical Machine Translation (WML)… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Webpages_Scraped_Translated.neptun.scraper
Data in this dataset
Docker & NPM
Scraped using crawl4ai.
The NPM and Docker data was scraped from docs.docker.com and docs.npmjs.com and processed using GPT-4 resulting in docker_documentation.jsonl and npm_documentation.jsonl.
The file training-data-v1.jsonl also includes Titanium, dockerNLcommands and docker_ps.
GitHub
Scraped using firecrawl.
The GitHub data was scraped from docs.github.com/en using firecrawl.A few pages might be missing in the… See the full description on the dataset page: https://huggingface.co/datasets/neptun-org/neptun.scraper.greenhouse-jobs-scraper
Greenhouse Jobs Scraper
Scrape every public job posting from any Greenhouse company job board: title, department, location, remote flag, seniority, advertised salary, full description and apply URL.
Rows in this dataset
14,091
Fields
36
Collector runs behind it
61
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/greenhouse-jobs-scraper/ — 92 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/greenhouse-jobs-scraper.Protocol_Scrape
Dataset Card for "Protocol_Scrape"
More Information needed
scrapegraph-100k-finetuning
ScrapeGraphAI 100k finetuning
Dataset Summary
A finetuning-ready derivative of ScrapeGraphAI-100k: schema-constrained web extraction examples where a model must produce JSON conforming to a user-defined JSON schema given Markdown-converted page content.
Split
Rows
Targets
train
25,244
GPT-5-nano regenerated targets
test
2,808
GPT-5-nano regenerated targets
human_eval
100
Human labeled extractions (evaluation only)
Important: train/test… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraph-100k-finetuning.arena.ai-code-leaderboard-scrapedlink: https://arena.ai/leaderboard/code/
steam-game-reviews-scraper-sample-data
Steam Game & Reviews Scraper
Scrape Steam game metadata, pricing, genres, Metacritic scores & user reviews using Steam's public API. Supports bulk app IDs, store URLs & keyword search. No proxy needed.
What the actor scrapes
Steam Game & Reviews Scraper — Steam Store Data & User Reviews to JSON/CSV Scrape game metadata and user reviews from the Steam Store using Steam's public JSON API. This Steam scraper extracts prices, discounts, genres, Metacritic scores… See the full description on the dataset page: https://huggingface.co/datasets/logiover/steam-game-reviews-scraper-sample-data.scrape-content-dataset-v1
Scrape Content Dataset v1
A human-curated benchmark dataset for evaluating web scraping engines on content quality.
Overview
This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time.
Dataset Structure
CSV format with columns:
id: Sequential identifier
url:… See the full description on the dataset page: https://huggingface.co/datasets/firecrawl/scrape-content-dataset-v1.scraped-hotel-reviewsgrants-gov-scraper
Grants.gov Scraper · Grant Opportunities, Agencies & Awards
Scrape US federal grant opportunities, funding announcements, and agency award notices from Grants.gov by keyword, agency, category, eligibility, and status.
Rows in this dataset
1,488
Fields
12
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/grants-gov-scraper.nowiki_second_scrape_merged
Dataset Card for "nowiki_second_scrape_merged"
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/jkorsvik/nowiki_second_scrape_merged.ekpatagolpo-scrape-bangla-literature
Ekpatagolpo Bengali Stories Archive
Request More ScrapesOrder Private Scrapes
Overview
This repository contains a large-scale, curated text dataset scraped from ekpatagolpo.com. The primary goal of this archive is to preserve a massive collection of purely human-written Bengali literature and stories (Bangla Golpo), creating a distinct record of human creativity separate from AI-generated text.
Purpose and Usage
This dataset is published publicly under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/ekpatagolpo-scrape-bangla-literature.defillama-protocols-scraper-sample-data
DefiLlama Protocols Scraper
Scrape all 7,000+ DeFi protocols from DefiLlama in one run — TVL, 1h/1d/7d TVL change, market cap, category, chains and links. Filter by chain, category and TVL. Schedule it daily to track the entire DeFi landscape.
What the actor scrapes
🦙 DefiLlama Protocols Scraper — Scrape All DeFi Protocols & TVL Data Scrape all 7,000+ DeFi protocols from DefiLlama in a single run and export them to JSON, CSV or Excel. This DefiLlama scraper… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-protocols-scraper-sample-data.shopify-store-products-scraper
Shopify Store Scraper
Scrape products, prices, discounts, variants and stock from any Shopify store's public product JSON. No login, no API key, no headless browser.
Rows in this dataset
11,720
Fields
33
Collector runs behind it
72
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/shopify-store-products-scraper/ — 1,271 entity pages
Run the collector yourself
https://apify.com/reapx/shopify-store-products-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-store-products-scraper.boardgamegeek-scraper
BoardGameGeek Scraper · Games, Ratings, Designers & Mechanics
Scrape board games, release years, player counts, categories, mechanics, designers, artists, and publishers from BoardGameGeek. HTTP only, pay-per-event pricing.
Rows in this dataset
437
Fields
21
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/boardgamegeek-scraper.
