datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AoPS-Scrape
AoPS-Scrape
Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints.
Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session.
Splits
Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split:
Split
Rows
Notes
deduplicated
29,964
One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.tool-scraper_daniel_20260827_124546This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/tool-scraper_daniel_20260827_124546.DrugHub-scrape
DrugHub Market Snapshot, September 2026
A complete, text-only capture of the public listing, vendor, and review pages of
DrugHub, a Monero-only darknet market operating since 2023. Everything here was
visible to any visitor without an account. Doesn't include any images.
Collected 16-17 September 2026. Enriched with model-derived labels
(typesafe/jev-1.13) on 19 September 2026; see the listing_enrichment
table and the Enrichment section below.
What's in it… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/DrugHub-scrape.scrape_residueThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue.steam-game-reviews-scraper-sample-data
Steam Game & Reviews Scraper
Scrape Steam game metadata, pricing, genres, Metacritic scores & user reviews using Steam's public API. Supports bulk app IDs, store URLs & keyword search. No proxy needed.
What the actor scrapes
Steam Game & Reviews Scraper — Steam Store Data & User Reviews to JSON/CSV Scrape game metadata and user reviews from the Steam Store using Steam's public JSON API. This Steam scraper extracts prices, discounts, genres, Metacritic scores… See the full description on the dataset page: https://huggingface.co/datasets/logiover/steam-game-reviews-scraper-sample-data.scrape_residue_20260916_110935This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue_20260916_110935.nowiki_second_scrape_merged
Dataset Card for "nowiki_second_scrape_merged"
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/jkorsvik/nowiki_second_scrape_merged.defillama-protocols-scraper-sample-data
DefiLlama Protocols Scraper
Scrape all 7,000+ DeFi protocols from DefiLlama in one run — TVL, 1h/1d/7d TVL change, market cap, category, chains and links. Filter by chain, category and TVL. Schedule it daily to track the entire DeFi landscape.
What the actor scrapes
🦙 DefiLlama Protocols Scraper — Scrape All DeFi Protocols & TVL Data Scrape all 7,000+ DeFi protocols from DefiLlama in a single run and export them to JSON, CSV or Excel. This DefiLlama scraper… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-protocols-scraper-sample-data.ScrapedJobslinkedin-top-content-scraper-sample-data
LinkedIn Top Content & Top Voices Scraper
Scrapes LinkedIn's public Top Content directory to extract curated high-engagement posts and Top Voice influencers across 40+ categories. Get post text, author profiles, follower counts, reaction metrics, and Top Voice badges. No login, no cookies, no account ban risk. $2 per 1,000 posts.
What the actor scrapes
LinkedIn Top Content & Top Voices Scraper Scrape LinkedIn's public Top Content directory — a curated archive of… See the full description on the dataset page: https://huggingface.co/datasets/logiover/linkedin-top-content-scraper-sample-data.hacker-news-scraped-storiesusaspending-gov-scraper-sample-data
USASpending.gov Federal Awards Scraper
Scrape US federal contracts, grants and awards from the official USASpending.gov API — no login, no API key, no blocking. Award ID, recipient, amount, agency, dates and place of performance. Filter by type, date and keyword. Hundreds of thousands of awards per run.
What the actor scrapes
🏛️ USASpending.gov Federal Awards Scraper — US Contracts, Grants & Awards to JSON & CSV Scrape US federal contracts, grants, loans and… See the full description on the dataset page: https://huggingface.co/datasets/logiover/usaspending-gov-scraper-sample-data.defillama-yields-scraper-sample-data
DefiLlama Yields Scraper
Scrape DeFi yield & APY pools from DefiLlama — APY, TVL, base/reward yield, 1d/7d/30d APY trend, impermanent-loss risk and volume for 20,000+ pools across every chain. Filter by chain, protocol, TVL and APY. Schedule it daily to track the best yields.
What the actor scrapes
💰 DefiLlama Yields Scraper — DeFi APY & TVL Pool Data Across All Chains Scrape DeFi yield and APY pools from DefiLlama, the most trusted DeFi data source. This Apify… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-yields-scraper-sample-data.lagou-tech-jobs-scraper-sample-data
Lagou Tech Jobs Scraper (拉勾网)
Extract thousands of tech job listings from Lagou.com (拉勾网), China's largest IT recruitment platform. Scrape salary ranges, tech stacks, company details, funding stages, and more from ByteDance, Alibaba, Tencent, Baidu, and 100,000+ Chinese tech companies. No browser needed — fast, cheap, scalable.
What the actor scrapes
Lagou Tech Jobs Scraper (拉勾网) — Scrape China Tech Jobs, Salaries & Company Data Scrape Lagou.com (拉勾网), China's #1… See the full description on the dataset page: https://huggingface.co/datasets/logiover/lagou-tech-jobs-scraper-sample-data.google-ads-transparency-scraper-sample-data
Google Ads Transparency Center Scraper
Scrapes every Google ad your competitors run — Search, Display, Shopping, YouTube, Maps. Multi-domain batch, multi-region, with optional impressions and spend enrichment. No login required.
What the actor scrapes
🎯 Google Ads Transparency Center Scraper — Competitor Ads, Impressions & Spend Scrape the Google Ads Transparency Centerat scale and extract every Google ad your competitors are running across Search, Display… See the full description on the dataset page: https://huggingface.co/datasets/logiover/google-ads-transparency-scraper-sample-data.openstreetmap-business-poi-scraper-sample-data
OpenStreetMap Business & POI Scraper
Scrape businesses and points of interest from OpenStreetMap via Overpass API. Extract name, address, phone, website, opening hours and GPS coordinates for any city worldwide. Free alternative to Google Maps API. No API key needed.
What the actor scrapes
🗺️ OpenStreetMap Business & POI Scraper — Scrape Businesses & Points of Interest, No API Key Scrape businesses and points of interest from OpenStreetMap using the free… See the full description on the dataset page: https://huggingface.co/datasets/logiover/openstreetmap-business-poi-scraper-sample-data.hacker-news-scraped-stories-filteredkalkibaat1-XOkZRgip-scraped-data-Final-Evalkalkibaat1-ZamcvGj3-scraped-data-Final-Evalsec-edgar-form-d-scraper-sample-data
SEC EDGAR Form D Scraper - Startup Funding Leads
Scrape SEC EDGAR Form D filings to find recently funded startups. Extract company name, funding amount, CEO/director names, contact info, industry & more. Perfect for B2B sales prospecting & investor research.
What the actor scrapes
SEC EDGAR Form D Scraper — Find Funded Startups & B2B Leads Scrape SEC EDGAR Form D (Regulation D) filings and turn public startup-funding disclosures into structured, actionable lead… See the full description on the dataset page: https://huggingface.co/datasets/logiover/sec-edgar-form-d-scraper-sample-data.Hini_Wiki_Scraped_Arkascraped_amazon__datascraped-forum-postsgreenhouse-job-board-scraper-sample-data
Greenhouse Job Board API — Jobs, Departments & Offices
Unofficial Greenhouse Job Board API in one Apify actor. Scrape jobs, full descriptions, departments and offices from any company on Greenhouse — Airbnb, Stripe, Anthropic, Mistral AI, Doctolib, Datadog, Notion. Pure HTTP, no auth, parallel batch. For HR tech, ATS, lead gen and AI agents.
What the actor scrapes
🌱 Greenhouse Job Board API — Scrape Tech Jobs, Departments & Offices The unofficial Greenhouse Job… See the full description on the dataset page: https://huggingface.co/datasets/logiover/greenhouse-job-board-scraper-sample-data.product-hunt-daily-launches-scraper-sample-data
Product Hunt Daily Launches Scraper
Scrape Product Hunt daily launches with votes, topics, descriptions, screenshots and video links. Supports daily snapshots, date ranges and topic filtering. No API key needed — works out of the box.
What the actor scrapes
Product Hunt Daily Launches Scraper Scrape Product Hunt daily launches, votes, maker profiles, topics, comments, and rankings via the official Product Hunt GraphQL API v2. Supports today's launches, custom… See the full description on the dataset page: https://huggingface.co/datasets/logiover/product-hunt-daily-launches-scraper-sample-data.habitaclia-com-spain-scraper-sample-data
Habitaclia.com Spain Scraper
Scrape real estate listings from Habitaclia.com — leading Spanish property portal owned by Adevinta, with deep coverage in Catalonia and major cities. Extract apartments, houses, attics, duplexes, offices, land and parking by region or city, with price (EUR), area (m²), rooms, location and images.
What the actor scrapes
Habitaclia.com Spain Property Scraper — Spanish Real Estate Listings to JSON/CSV/Excel Scrape real estate listings… See the full description on the dataset page: https://huggingface.co/datasets/logiover/habitaclia-com-spain-scraper-sample-data.pisos-com-property-scraper-sample-data
Pisos.com Property Scraper
Scrape real estate listings from Pisos.com — Spain's #3 property portal, owned by Vocento. Extract apartments, houses, attics, duplexes, studios, lofts, offices, garages and storage (sale, rent or new build) by Spanish region, province or city, with price (EUR), area (m²), rooms, bathrooms and GPS.
What the actor scrapes
🏠 Pisos.com Property Scraper — Scrape Spain Real Estate Listings & Prices Scrape real estate listings from Pisos.com… See the full description on the dataset page: https://huggingface.co/datasets/logiover/pisos-com-property-scraper-sample-data.scrapegraphai-100k
ScrapeGraphAI 100k
Dataset Summary
This dataset contains 100k curated structured output examples from the ScrapeGraphAI OSS library logs collected in Q2-Q3 of 2025. Each example captures an LLM's attempt to extract structured data from web content following a user-defined JSON schema.
The dataset was derived from 9 million raw PostHog events, processed to extract schema complexity metrics, and balanced to ensure diversity across unique schemas.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Rendy45/scrapegraphai-100k.pick_scraper_from_rackThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 25000,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lyl472324464/pick_scraper_from_rack.lagou-tech-jobs-scraper-sample-data
Lagou Tech Jobs Scraper (拉勾网)
Extract thousands of tech job listings from Lagou.com (拉勾网), China's largest IT recruitment platform. Scrape salary ranges, tech stacks, company details, funding stages, and more from ByteDance, Alibaba, Tencent, Baidu, and 100,000+ Chinese tech companies. No browser needed — fast, cheap, scalable.
What the actor scrapes
Lagou Tech Jobs Scraper (拉勾网) — Scrape China Tech Jobs, Salaries & Company Data Scrape Lagou.com (拉勾网)… See the full description on the dataset page: https://huggingface.co/datasets/gansusu/lagou-tech-jobs-scraper-sample-data.
