CoolFace
Datasetpublic

reapxdev/hackernews-scraper

Hacker News Scraper · Stories, Comments, Points & Domains Scrape Hacker News stories, comments, Ask HN, Show HN, point thresholds, date ranges, and linked web domains via the official HN Search API by Algolia with rich search filters and domain extraction. Rows in this dataset 2,375 Fields 15 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/hackernews-scraper.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes40downloads
Dataset Card

reapX — public sources in, addressable records out

Hacker News Scraper · Stories, Comments, Points & Domains

Scrape Hacker News stories, comments, Ask HN, Show HN, point thresholds, date ranges, and linked web domains via the official HN Search API by Algolia with rich search filters and domain extraction.

Rows in this dataset2,375
Fields15
Collector runs behind it50
Most recent observation2026-08-03

What this is

Every row here was returned by a real run of a public collector. Nothing is generated from a template over a keyword list: a row exists because a run observed it.

Provenance

Each row carries _run_id and _dataset_id, naming the collector run that produced it, so any row can be traced back to the run that observed it. Rows observed by more than one run are deduplicated on content; 77 duplicate observations were collapsed.

Files

  • hackernews-scraper.jsonl — one JSON object per row, the canonical form
  • hackernews-scraper.csv — the same rows flattened; nested values are JSON-encoded within their cell so they round-trip
  • dataset.json — schema.org Dataset metadata

Loading it

python
from datasets import load_dataset
ds = load_dataset("reapxdev/hackernews-scraper", split="train")

Related

  • All sources: <https://reapx.dev/data/> · machine-readable index: <https://reapx.dev/llms.txt>
  • The collector is a public Apify Actor; agents reach it through <https://mcp.apify.com>

Licence

Collected from public sources. This metadata and the published pages are CC BY 4.0; the underlying records remain under the terms of their originating source.