datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hackernews-comments
Hackernews Comments Dataset
A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape
Dataset contents
No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload:
{
"by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.hacker-news
Hacker News Dataset
Dataset Summary
This dataset is derived from the official Hacker News data provided via the Hacker News Firebase API. It contains user-generated content including stories, comments, and metadata from the Hacker News platform.
Hacker News is a social news website focusing on computer science, entrepreneurship, and technology. The dataset captures real-world discussions, technical conversations, and community interactions over time.
The data was… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/hacker-news.hackernews-stories
A HackerNews Stories dataset
This dataset is based on nixiesearch/hackernews-comments dataset:
for each item of type=story we downloaded the target URL. Out of ~3.8M stories ~2.1M are still reachable.
each story HTML was parsed using trafilatura library
we store article text in markdown format along with all page-specific metadata.
Dataset stats
date coverage: xx.2006-09.2024, same as in upstream nixiesearch/hackernews-comments dataset
total scraped pages: 2150271… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-stories.hackernews-scraper
Hacker News Scraper · Stories, Comments, Points & Domains
Scrape Hacker News stories, comments, Ask HN, Show HN, point thresholds, date ranges, and linked web domains via the official HN Search API by Algolia with rich search filters and domain extraction.
Rows in this dataset
2,375
Fields
15
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/hackernews-scraper.pile_hackernewsthe-pile-hackernews-refined-by-data-juicer
The Pile -- HackerNews (refined by Data-Juicer)
A refined version of HackerNews dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.8G).
Dataset Information
Number of samples: 371,331 (Keep ~99.55% from the original dataset)
Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-hackernews-refined-by-data-juicer.hackernewshackernews-training
