datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hacker-news-rss
Hacker News RSS Feed Directory
TL;DR — We visited every unique domain ever posted to Hacker News, found
which ones publish RSS/Atom feeds, and packaged the results as monthly
parquet snapshots with rich metadata.
623,957 feeds discovered across 1,755,955 hosts,
spanning 232 months from 2006-10 to 2026-03.
Last updated: 2026-04-05T09:21:39Z
Why this exists
RSS is not dead — it's just hard to discover. The <link rel="alternate">
tag that points to a site's feed is… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news-rss.hacker-news-posts
Hacker News Stories Dataset
This is a dataset containing approximately 4 million stories from Hacker News, exported to a Parquet file. The dataset includes the following fields:
id (int64): The unique identifier of the story.
title (string): The title of the story.
url (string): The URL of the story.
score (int64): The score of the story.
time (int64): The time the story was posted, in Unix time.
comments (int64): The number of comments on the story.
author (string): The… See the full description on the dataset page: https://huggingface.co/datasets/julien040/hacker-news-posts.hacker-news
Hacker News posts and comments
This is a dataset of all HN posts and comments, current as of November 1, 2023.
hackernews-vector-search-datasetThe Hacker News dataset contains 28.74 million postings and their vector embeddings. The embeddings were generated using SentenceTransformers model all-MiniLM-L6-v2. The dimension of each embedding vector is 384.
Created by clickhouse more info: https://clickhouse.com/docs/getting-started/example-datasets/hackernews-vector-search-dataset
hacker-news-who-is-hiring-posts
Context
This dataset contains all first-level comments to Hacker News Who Is Hiring posts from April 2011 in various formats. All data is derived from the official Firebase API and no data cleansing has occurred with the exception for removing SEEKING FREELANCER from the start of such comments..
Who wants to be hired? and Seeking Freelancer posts are included. For privacy reasons, job seeker posts will not be included. Although the data is public, do not want to create an easily… See the full description on the dataset page: https://huggingface.co/datasets/brusic/hacker-news-who-is-hiring-posts.hacker-news-corpus-2007-2022
Hacker News corpus, 2007-Nov 2022
Dataset Description
Dataset Summary
Dataset Name: Hacker News Full Corpus (2007 - November 2022)
Description:
NOTE: I am not affiliated with Y Combinator.
This dataset is a July 2023 snapshot of YCombinator's BigQuery dump of the entire archive of posts and comments made on Hacker News. It contains posts from Hacker News' inception in 2007 through to November 16, 2022, when the BigQuery database was last updated.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/jkeisling/hacker-news-corpus-2007-2022.hackernews_hiring_postsThis dataset contains postings and comments from the following recurring threads on Hacker News
Ask HN: Who is hiring?
Ask HN: Who wants to be hired?
Freelancer? Seeking freelancer?
These post types are stored in datasets called hiring, wants_to_be_hired and freelancer respectively.
Each type of posting has occurred on a regular basis for several years. You can identify when each comment/listing was added through the CommentTime field. The ParentTitle also indicates the date of the parent… See the full description on the dataset page: https://huggingface.co/datasets/dansbecker/hackernews_hiring_posts.hacker-newsThis repository contains the datasets for hacker news, used by https://github.com/anantn/hn-chatgpt-plugin
As of June 2025, these are now exported as parquet files instead of sqlite for space efficiency
hacker-news-dataset
Hacker News Dataset (2025)
Dataset Description
A comprehensive dataset of Hacker News content from 2025, containing stories, comments, users, and their relationships. This dataset enables deep analysis of technical discussions, trends, and community dynamics on one of the most influential technology forums.
Dataset Summary
Total Records: 38.4M+ across 10 tables
Stories: 287K+ submissions including links, Show HNs, Ask HNs
Comments: 2.5M+ discussion… See the full description on the dataset page: https://huggingface.co/datasets/typedef-ai/hacker-news-dataset.hackernews-stories
Dataset Card for "hackernews-stories"
More Information needed
Pile-HackerNews-0.5B-6K-opt
Dataset Card for "Pile-HackerNews-0.5B-6K-opt"
More Information needed
hacker-news-scraped-storiespile-hackernews
Dataset Creation Process
These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered.
Citations
If you use this dataset, please cite the original Pile papers:
@article{gao2020pile,
title={The Pile: An 800GB dataset of diverse text for language modeling},
author={Gao, Leo and Biderman, Stella and Black, Sid and Golding, Laurence and… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-hackernews.hacker-news-scraped-stories-filteredhacker-news-text-search
Hacker News text + substring patterns
Sampled comments and stories from the full year 2025 of the public
Hacker News archive, paired with
small curated dictionaries of substring patterns and precomputed
match labels. The intended use is testing text-search and
substring-matching code on real, messy English text: multi-byte
characters, HTML entities, embedded URLs, mixed casing, CVE
identifiers, version strings, and the long tail of forum slang.
Layout at a glance… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news-text-search.hacker_news_prompt_completion
Dataset Card for "hacker_news_prompt_completion"
More Information needed
hacker_news_top_comment
Dataset Card for "hacker_news_top_comment"
More Information needed
Pile-HackerNews-0.5B-8K-opt
Dataset Card for "Pile-HackerNews-0.5B-8K-opt"
More Information needed
hacker-news-discussion-summarization-large
Dataset Card for Hacker News Discussion Summarization - Large
Dataset Summary
This dataset comprises 14,531 records of Hacker News front-page stories collected over 516 days. Each record includes the story's metadata and its associated discussion threads, formatted to facilitate the development of summarization models.
Supported Tasks and Leaderboards
The primary task supported by this dataset is summarization, specifically targeting the summarization of… See the full description on the dataset page: https://huggingface.co/datasets/georgeck/hacker-news-discussion-summarization-large.hacker-news-regressor-datasethacker_news_texthackernewsTop 1000 HackerNews links for every month from Oct. 2006 to June 2025
hacker-news-who-is-hiring-scraper-sample-data
Hacker News Who Is Hiring Scraper – Jobs, Salary & Email
Scrape structured job listings from Hacker News 'Who is Hiring?' monthly threads. Extracts company, role, location, salary, remote policy and tech stack — no AI, no API key, no proxy needed.
What the actor scrapes
Hacker News Who Is Hiring Scraper — Jobs, Salary & Tech Stack Data Scrape structured job listings from Hacker News "Ask HN: Who is Hiring?"monthly threads. Extracts company name, role, location… See the full description on the dataset page: https://huggingface.co/datasets/logiover/hacker-news-who-is-hiring-scraper-sample-data.hacker-news-discussion-summarization-smallmimir-sage-paraphrased-hackernews
Overview
This dataset is derived from the upstream MIMIR benchmark dataset (iamgroot42/mimir) and augments each example with multiple paraphrased variants produced by SAGE/SAGER pipelines using different LLM backends (DeepSeek-V3.2, Gemini-2.5-Flash, Grok-4.1-Fast).
Source subset: HackerNews
Columns
member: member text
nonmember: non-member text
sage_deepseek: SAGE paraphrase using DeepSeek-V3.2
sage_gemini: SAGE paraphrase using Gemini-2.5-Flash
sage_grok: SAGE… See the full description on the dataset page: https://huggingface.co/datasets/bilgehanertan/mimir-sage-paraphrased-hackernews.hackernewsscaling_mia_the_pile_00_HackerNewshacker-news-discussion-summarizationhacker_newshackernews
hackernews
This dataset is produced and published automatically by DataMax. It contains the following assets:
most_frequent_words
top_stories
most_frequent_words
Get the top 25 most frequent words in the titles of the top 100 HackerNews stories.
This dataset is produced and published automatically by DataMax.
Dataset Statistics
Number of rows: 1
Number of columns: 25
Sample Data
hn
show
–
new
from
why
answer
api
using
birth… See the full description on the dataset page: https://huggingface.co/datasets/substrate-labs/hackernews.
