datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AoPS-Scrape
AoPS-Scrape
Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints.
Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session.
Splits
Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split:
Split
Rows
Notes
deduplicated
29,964
One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.tatoeba-nusax-scrape-mt-concatgleif-lei-scraper-sample-data
GLEIF LEI Scraper
Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run.
What the actor scrapes
🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.penetration_testing_scraped_dataset
Dataset Card for "penetration_testing_scraped_dataset"
More Information needed
tool-scraper_daniel_20260827_124546This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/tool-scraper_daniel_20260827_124546.fitcheck-scraped-multiviewfitcheck-scraped-v1DrugHub-scrape
DrugHub Market Snapshot, September 2026
A complete, text-only capture of the public listing, vendor, and review pages of
DrugHub, a Monero-only darknet market operating since 2023. Everything here was
visible to any visitor without an account. Doesn't include any images.
Collected 16-17 September 2026. Enriched with model-derived labels
(typesafe/jev-1.13) on 19 September 2026; see the listing_enrichment
table and the Enrichment section below.
What's in it… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/DrugHub-scrape.scrape_residueThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue.Protocol_Scrape
Dataset Card for "Protocol_Scrape"
More Information needed
English_French_Webpages_Scraped_Translated
English French Webpages Scraped Translated
Dataset Summary
French/English parallel texts for training translation models. Over 17.1 million sentences in French and English. Dataset created by Chris Callison-Burch, who crawled millions of web pages and then used a set of simple heuristics to transform French URLs onto English URLs, and assumed that these documents are translations of each other. This is the main dataset of Workshop on Statistical Machine Translation (WML)… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Webpages_Scraped_Translated.scrapegraph-100k-finetuning
ScrapeGraphAI 100k finetuning
Dataset Summary
A finetuning-ready derivative of ScrapeGraphAI-100k: schema-constrained web extraction examples where a model must produce JSON conforming to a user-defined JSON schema given Markdown-converted page content.
Split
Rows
Targets
train
25,244
GPT-5-nano regenerated targets
test
2,808
GPT-5-nano regenerated targets
human_eval
100
Human labeled extractions (evaluation only)
Important: train/test… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraph-100k-finetuning.steam-game-reviews-scraper-sample-data
Steam Game & Reviews Scraper
Scrape Steam game metadata, pricing, genres, Metacritic scores & user reviews using Steam's public API. Supports bulk app IDs, store URLs & keyword search. No proxy needed.
What the actor scrapes
Steam Game & Reviews Scraper — Steam Store Data & User Reviews to JSON/CSV Scrape game metadata and user reviews from the Steam Store using Steam's public JSON API. This Steam scraper extracts prices, discounts, genres, Metacritic scores… See the full description on the dataset page: https://huggingface.co/datasets/logiover/steam-game-reviews-scraper-sample-data.scrape_residue_20260916_110935This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue_20260916_110935.defillama-protocols-scraper-sample-data
DefiLlama Protocols Scraper
Scrape all 7,000+ DeFi protocols from DefiLlama in one run — TVL, 1h/1d/7d TVL change, market cap, category, chains and links. Filter by chain, category and TVL. Schedule it daily to track the entire DeFi landscape.
What the actor scrapes
🦙 DefiLlama Protocols Scraper — Scrape All DeFi Protocols & TVL Data Scrape all 7,000+ DeFi protocols from DefiLlama in a single run and export them to JSON, CSV or Excel. This DefiLlama scraper… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-protocols-scraper-sample-data.nowiki_second_scrape_merged
Dataset Card for "nowiki_second_scrape_merged"
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/jkorsvik/nowiki_second_scrape_merged.India-Lok-Sabha-Debates-Dataset-ScraperScrapedJobslinkedin-top-content-scraper-sample-data
LinkedIn Top Content & Top Voices Scraper
Scrapes LinkedIn's public Top Content directory to extract curated high-engagement posts and Top Voice influencers across 40+ categories. Get post text, author profiles, follower counts, reaction metrics, and Top Voice badges. No login, no cookies, no account ban risk. $2 per 1,000 posts.
What the actor scrapes
LinkedIn Top Content & Top Voices Scraper Scrape LinkedIn's public Top Content directory — a curated archive of… See the full description on the dataset page: https://huggingface.co/datasets/logiover/linkedin-top-content-scraper-sample-data.ekpatagolpo-scrape-bangla-literature
Ekpatagolpo Bengali Stories Archive
Request More ScrapesOrder Private Scrapes
Overview
This repository contains a large-scale, curated text dataset scraped from ekpatagolpo.com. The primary goal of this archive is to preserve a massive collection of purely human-written Bengali literature and stories (Bangla Golpo), creating a distinct record of human creativity separate from AI-generated text.
Purpose and Usage
This dataset is published publicly under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/ekpatagolpo-scrape-bangla-literature.web-scraper-datasetPolyMath-Scraped-Raw
PolyMath Scraped
PolyMath is a curated dataset of 11,090 high-difficulty mathematical problems designed for training reasoning models. Built for the AIMO Math Corpus Prize. Existing math datasets (NuminaMath-1.5, OpenMathReasoning) suffer from high noise rates in their hardest samples and largely unusable proof-based problems.
PolyMath addresses both issues through:
Data scraping: problems sourced from official competition PDFs absent from popular datasets, using a… See the full description on the dataset page: https://huggingface.co/datasets/AIMO-Corpus/PolyMath-Scraped-Raw.NYT-Connections-Verl-Scrapedhacker-news-scraped-storiesscraped-episodesusaspending-gov-scraper-sample-data
USASpending.gov Federal Awards Scraper
Scrape US federal contracts, grants and awards from the official USASpending.gov API — no login, no API key, no blocking. Award ID, recipient, amount, agency, dates and place of performance. Filter by type, date and keyword. Hundreds of thousands of awards per run.
What the actor scrapes
🏛️ USASpending.gov Federal Awards Scraper — US Contracts, Grants & Awards to JSON & CSV Scrape US federal contracts, grants, loans and… See the full description on the dataset page: https://huggingface.co/datasets/logiover/usaspending-gov-scraper-sample-data.defillama-yields-scraper-sample-data
DefiLlama Yields Scraper
Scrape DeFi yield & APY pools from DefiLlama — APY, TVL, base/reward yield, 1d/7d/30d APY trend, impermanent-loss risk and volume for 20,000+ pools across every chain. Filter by chain, protocol, TVL and APY. Schedule it daily to track the best yields.
What the actor scrapes
💰 DefiLlama Yields Scraper — DeFi APY & TVL Pool Data Across All Chains Scrape DeFi yield and APY pools from DefiLlama, the most trusted DeFi data source. This Apify… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-yields-scraper-sample-data.scraped_xsum1jobicy-remote-jobs-scraper-sample-data
Jobicy Remote Jobs Scraper
Scrape remote job listings from Jobicy. Filter by keyword, industry, job type, level and geography. Run on a schedule with the posted-since filter to capture only new jobs.
What the actor scrapes
💼 Jobicy Remote Jobs Scraper — Scrape Remote Job Listings & Export to JSON/CSV Scrape remote job listings from Jobicy — one of the most popular remote-work job boards — straight from its public API. This Jobicy scraper delivers a clean… See the full description on the dataset page: https://huggingface.co/datasets/logiover/jobicy-remote-jobs-scraper-sample-data.scraped-forum-threads
