datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tool-scraper_daniel_20260827_124546This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/tool-scraper_daniel_20260827_124546.scrape_residueThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue.greenhouse-jobs-scraper
Greenhouse Jobs Scraper
Scrape every public job posting from any Greenhouse company job board: title, department, location, remote flag, seniority, advertised salary, full description and apply URL.
Rows in this dataset
14,091
Fields
36
Collector runs behind it
61
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/greenhouse-jobs-scraper/ — 92 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/greenhouse-jobs-scraper.steam-game-reviews-scraper-sample-data
Steam Game & Reviews Scraper
Scrape Steam game metadata, pricing, genres, Metacritic scores & user reviews using Steam's public API. Supports bulk app IDs, store URLs & keyword search. No proxy needed.
What the actor scrapes
Steam Game & Reviews Scraper — Steam Store Data & User Reviews to JSON/CSV Scrape game metadata and user reviews from the Steam Store using Steam's public JSON API. This Steam scraper extracts prices, discounts, genres, Metacritic scores… See the full description on the dataset page: https://huggingface.co/datasets/logiover/steam-game-reviews-scraper-sample-data.scrape_residue_20260916_110935This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/scrape_residue_20260916_110935.defillama-protocols-scraper-sample-data
DefiLlama Protocols Scraper
Scrape all 7,000+ DeFi protocols from DefiLlama in one run — TVL, 1h/1d/7d TVL change, market cap, category, chains and links. Filter by chain, category and TVL. Schedule it daily to track the entire DeFi landscape.
What the actor scrapes
🦙 DefiLlama Protocols Scraper — Scrape All DeFi Protocols & TVL Data Scrape all 7,000+ DeFi protocols from DefiLlama in a single run and export them to JSON, CSV or Excel. This DefiLlama scraper… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-protocols-scraper-sample-data.boardgamegeek-scraper
BoardGameGeek Scraper · Games, Ratings, Designers & Mechanics
Scrape board games, release years, player counts, categories, mechanics, designers, artists, and publishers from BoardGameGeek. HTTP only, pay-per-event pricing.
Rows in this dataset
437
Fields
21
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/boardgamegeek-scraper.shopify-store-products-scraper
Shopify Store Scraper
Scrape products, prices, discounts, variants and stock from any Shopify store's public product JSON. No login, no API key, no headless browser.
Rows in this dataset
11,720
Fields
33
Collector runs behind it
72
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/shopify-store-products-scraper/ — 1,271 entity pages
Run the collector yourself
https://apify.com/reapx/shopify-store-products-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-store-products-scraper.linkedin-top-content-scraper-sample-data
LinkedIn Top Content & Top Voices Scraper
Scrapes LinkedIn's public Top Content directory to extract curated high-engagement posts and Top Voice influencers across 40+ categories. Get post text, author profiles, follower counts, reaction metrics, and Top Voice badges. No login, no cookies, no account ban risk. $2 per 1,000 posts.
What the actor scrapes
LinkedIn Top Content & Top Voices Scraper Scrape LinkedIn's public Top Content directory — a curated archive of… See the full description on the dataset page: https://huggingface.co/datasets/logiover/linkedin-top-content-scraper-sample-data.shopify-app-store-scraper
Shopify App Store Scraper
List Shopify App Store apps and get one structured row per app, with developer, star rating, review count, every advertised pricing plan and the app's rank in its category.
Rows in this dataset
3,880
Fields
26
Collector runs behind it
87
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/shopify-app-store-scraper/ — 1,625 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-app-store-scraper.stackoverflow-scraper
StackOverflow Scraper
Scrape Stack Overflow questions, answers, tags and user profiles through the public Stack Exchange API. Filter by tag, score, date, accepted status and full-text search. No login, no browser.
Rows in this dataset
16,719
Fields
47
Collector runs behind it
57
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/stackoverflow-scraper/ — 9,841 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/stackoverflow-scraper.atomic-scraper-leads
Atomic Scraper Leads
Public business listings scraped from Google Maps for leads.benjaminboyce.com.
This dataset is an export of the leads table from the Atomic / gmaps-scraper-suite pipeline. Each row is a local business listing with contact and location fields plus scrape metadata.
Freshness
Use scraped_at (UTC timestamps) as the freshness signal. This snapshot was exported on 2026-08-25.
Oldest scraped_at: 2026-08-09
Newest scraped_at: 2026-08-25
Listings… See the full description on the dataset page: https://huggingface.co/datasets/bensblueprints/atomic-scraper-leads.app-store-reviews-scraper
App Store Reviews Scraper
Scrape Apple App Store reviews, star ratings and app version history for any iOS app in any country storefront. No login, no API key.
Rows in this dataset
23,048
Fields
43
Collector runs behind it
88
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/app-store-reviews-scraper/ — 176 entity pages
Run the collector yourself
https://apify.com/reapx/app-store-reviews-scraper
What this is… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/app-store-reviews-scraper.kalshi-scraper
Kalshi Scraper · Event Contracts, Markets, Prices & Volume
Scrape Kalshi prediction markets, event contracts, option pricing, order book quotes, trading volume, open interest, and resolution rules. Export structured JSON, CSV, or Excel data.
Rows in this dataset
7,649
Fields
34
Collector runs behind it
49
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/kalshi-scraper.defillama-yields-scraper-sample-data
DefiLlama Yields Scraper
Scrape DeFi yield & APY pools from DefiLlama — APY, TVL, base/reward yield, 1d/7d/30d APY trend, impermanent-loss risk and volume for 20,000+ pools across every chain. Filter by chain, protocol, TVL and APY. Schedule it daily to track the best yields.
What the actor scrapes
💰 DefiLlama Yields Scraper — DeFi APY & TVL Pool Data Across All Chains Scrape DeFi yield and APY pools from DefiLlama, the most trusted DeFi data source. This Apify… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-yields-scraper-sample-data.steam-reviews-scraper
Steam Reviews Scraper · Game Reviews, Ratings & Playtime
Scrape Steam game reviews, ratings, playtime, helpfulness votes, and purchase types across any Steam App ID. Fast HTTP scraper, no login required.
Rows in this dataset
9,260
Fields
26
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a keyword list: a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/steam-reviews-scraper.semantic-scholar-scraper
Semantic Scholar Scraper · Papers, Authors, Citations & Venues
Scrape academic research papers, authors, citations, venues, and open-access metadata from Semantic Scholar API. Features rate-limit backoff resilience and pay-per-event pricing.
Rows in this dataset
450
Fields
19
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/semantic-scholar-scraper.agent-scraper-price-index-2026-09
Agent scraper price index, September 2026
The rows behind the report Agent scraper price index, September 2026 on cracked.ai: what 1,000 results cost for 18 of the most requested scraping and search capabilities, tool by tool, through three routes.
Method
For each capability, every live candidate tool on Cracked is listed with three prices for 1,000 results: the price billed on Cracked's own provider account (cost_1000_cracked_usd, provider price plus the $0.001… See the full description on the dataset page: https://huggingface.co/datasets/crackedvibe/agent-scraper-price-index-2026-09.wordpress-plugins-scraper
WordPress Plugins Scraper · Plugins, Installs & Ratings
Scrape WordPress plugins directory by tags, search queries, author accounts, install bands, and rating filters. Extract ratings, active installs, tags, release details, and author links.
Rows in this dataset
1,660
Fields
23
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/wordpress-plugins-scraper.arxiv-papers-scraper
arXiv Papers Scraper
Search arXiv and export papers with full abstracts, author lists, subject categories, DOIs, journal references and PDF links. Filter by subject class, keyword, author, affiliation or date window.
Rows in this dataset
21,722
Fields
26
Collector runs behind it
92
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/arxiv-papers-scraper/ — 10,624 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/arxiv-papers-scraper.remoteok-jobs-scraper
RemoteOK Jobs Scraper
Search RemoteOK and get one structured row per remote job, with company, role, tags, posted date, published salary range where given, and a direct apply link.
Rows in this dataset
3,001
Fields
30
Collector runs behind it
53
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/remoteok-jobs-scraper/ — 1,108 entity pages
Run the collector yourself
https://apify.com/reapx/remoteok-jobs-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/remoteok-jobs-scraper.github-repo-scraper
GitHub Repo Scraper · Repositories, Stars, Topics & Languages
Scrape GitHub repositories by language, topic, star count, license, organization, and pushed date window. Returns clean structured repo metrics and metadata without authentication.
Rows in this dataset
2,481
Fields
27
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/github-repo-scraper/ — 2,071 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/github-repo-scraper.zenodo-scraper
Zenodo Scraper · Research Records, DOIs, Authors & Files
Scrape open research records, DOIs, publications, datasets, software, authors, and file metadata from Zenodo. Fast HTTP scraper charging per returned record with tiered pricing.
Rows in this dataset
1,965
Fields
27
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/zenodo-scraper.lagou-tech-jobs-scraper-sample-data
Lagou Tech Jobs Scraper (拉勾网)
Extract thousands of tech job listings from Lagou.com (拉勾网), China's largest IT recruitment platform. Scrape salary ranges, tech stacks, company details, funding stages, and more from ByteDance, Alibaba, Tencent, Baidu, and 100,000+ Chinese tech companies. No browser needed — fast, cheap, scalable.
What the actor scrapes
Lagou Tech Jobs Scraper (拉勾网) — Scrape China Tech Jobs, Salaries & Company Data Scrape Lagou.com (拉勾网), China's #1… See the full description on the dataset page: https://huggingface.co/datasets/logiover/lagou-tech-jobs-scraper-sample-data.hackernews-scraper
Hacker News Scraper · Stories, Comments, Points & Domains
Scrape Hacker News stories, comments, Ask HN, Show HN, point thresholds, date ranges, and linked web domains via the official HN Search API by Algolia with rich search filters and domain extraction.
Rows in this dataset
2,375
Fields
15
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/hackernews-scraper.google-ads-transparency-scraper-sample-data
Google Ads Transparency Center Scraper
Scrapes every Google ad your competitors run — Search, Display, Shopping, YouTube, Maps. Multi-domain batch, multi-region, with optional impressions and spend enrichment. No login required.
What the actor scrapes
🎯 Google Ads Transparency Center Scraper — Competitor Ads, Impressions & Spend Scrape the Google Ads Transparency Centerat scale and extract every Google ad your competitors are running across Search, Display… See the full description on the dataset page: https://huggingface.co/datasets/logiover/google-ads-transparency-scraper-sample-data.openstreetmap-business-poi-scraper-sample-data
OpenStreetMap Business & POI Scraper
Scrape businesses and points of interest from OpenStreetMap via Overpass API. Extract name, address, phone, website, opening hours and GPS coordinates for any city worldwide. Free alternative to Google Maps API. No API key needed.
What the actor scrapes
🗺️ OpenStreetMap Business & POI Scraper — Scrape Businesses & Points of Interest, No API Key Scrape businesses and points of interest from OpenStreetMap using the free… See the full description on the dataset page: https://huggingface.co/datasets/logiover/openstreetmap-business-poi-scraper-sample-data.openalex-scraper
OpenAlex Scraper · Works, Authors, Institutions & Citations
Scrape scholarly works, papers, citations, authors, institutions, and open-access metadata from the OpenAlex API. Fast HTTP scraper charging per returned record with tiered pricing.
Rows in this dataset
2,387
Fields
18
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/openalex-scraper/ — 2,387 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/openalex-scraper.usaspending-gov-scraper-sample-data
USASpending.gov Federal Awards Scraper
Scrape US federal contracts, grants and awards from the official USASpending.gov API — no login, no API key, no blocking. Award ID, recipient, amount, agency, dates and place of performance. Filter by type, date and keyword. Hundreds of thousands of awards per run.
What the actor scrapes
🏛️ USASpending.gov Federal Awards Scraper — US Contracts, Grants & Awards to JSON & CSV Scrape US federal contracts, grants, loans and… See the full description on the dataset page: https://huggingface.co/datasets/logiover/usaspending-gov-scraper-sample-data.apple-podcasts-scraper
Apple Podcasts Scraper · Shows, Episodes, Genres & Rankings
Scrape Apple Podcasts catalog, shows, episodes, top charts, genres, and rankings. HTTP-only iTunes Search API scraper for audio analytics, podcast discovery, and media datasets.
Rows in this dataset
2,189
Fields
20
Collector runs behind it
51
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/apple-podcasts-scraper.
