datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hacker-news-rss
Hacker News RSS Feed Directory
TL;DR — We visited every unique domain ever posted to Hacker News, found
which ones publish RSS/Atom feeds, and packaged the results as monthly
parquet snapshots with rich metadata.
623,957 feeds discovered across 1,755,955 hosts,
spanning 232 months from 2006-10 to 2026-03.
Last updated: 2026-04-05T09:21:39Z
Why this exists
RSS is not dead — it's just hard to discover. The <link rel="alternate">
tag that points to a site's feed is… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news-rss.hacker-news-posts
Hacker News Stories Dataset
This is a dataset containing approximately 4 million stories from Hacker News, exported to a Parquet file. The dataset includes the following fields:
id (int64): The unique identifier of the story.
title (string): The title of the story.
url (string): The URL of the story.
score (int64): The score of the story.
time (int64): The time the story was posted, in Unix time.
comments (int64): The number of comments on the story.
author (string): The… See the full description on the dataset page: https://huggingface.co/datasets/julien040/hacker-news-posts.hacker-news
Hacker News posts and comments
This is a dataset of all HN posts and comments, current as of November 1, 2023.
hackernews-vector-search-datasetThe Hacker News dataset contains 28.74 million postings and their vector embeddings. The embeddings were generated using SentenceTransformers model all-MiniLM-L6-v2. The dimension of each embedding vector is 384.
Created by clickhouse more info: https://clickhouse.com/docs/getting-started/example-datasets/hackernews-vector-search-dataset
hacker-news-corpus-2007-2022
Hacker News corpus, 2007-Nov 2022
Dataset Description
Dataset Summary
Dataset Name: Hacker News Full Corpus (2007 - November 2022)
Description:
NOTE: I am not affiliated with Y Combinator.
This dataset is a July 2023 snapshot of YCombinator's BigQuery dump of the entire archive of posts and comments made on Hacker News. It contains posts from Hacker News' inception in 2007 through to November 16, 2022, when the BigQuery database was last updated.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/jkeisling/hacker-news-corpus-2007-2022.hacker-newsThis repository contains the datasets for hacker news, used by https://github.com/anantn/hn-chatgpt-plugin
As of June 2025, these are now exported as parquet files instead of sqlite for space efficiency
hacker-news-dataset
Hacker News Dataset (2025)
Dataset Description
A comprehensive dataset of Hacker News content from 2025, containing stories, comments, users, and their relationships. This dataset enables deep analysis of technical discussions, trends, and community dynamics on one of the most influential technology forums.
Dataset Summary
Total Records: 38.4M+ across 10 tables
Stories: 287K+ submissions including links, Show HNs, Ask HNs
Comments: 2.5M+ discussion… See the full description on the dataset page: https://huggingface.co/datasets/typedef-ai/hacker-news-dataset.hacker-news-scraped-storieshacker-news-scraped-stories-filteredhacker-news-text-search
Hacker News text + substring patterns
Sampled comments and stories from the full year 2025 of the public
Hacker News archive, paired with
small curated dictionaries of substring patterns and precomputed
match labels. The intended use is testing text-search and
substring-matching code on real, messy English text: multi-byte
characters, HTML entities, embedded URLs, mixed casing, CVE
identifiers, version strings, and the long tail of forum slang.
Layout at a glance… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news-text-search.hacker-news-regressor-datasethackernewsTop 1000 HackerNews links for every month from Oct. 2006 to June 2025
hackernewshackernews
hackernews
This dataset is produced and published automatically by DataMax. It contains the following assets:
most_frequent_words
top_stories
most_frequent_words
Get the top 25 most frequent words in the titles of the top 100 HackerNews stories.
This dataset is produced and published automatically by DataMax.
Dataset Statistics
Number of rows: 1
Number of columns: 25
Sample Data
hn
show
–
new
from
why
answer
api
using
birth… See the full description on the dataset page: https://huggingface.co/datasets/substrate-labs/hackernews.hacker-news-search-scraper-sample-data
Hacker News Search Scraper
Scrape Hacker News stories, comments and polls at scale via the Algolia API — title, author, URL, text, points, comment count and tags. Search by keyword, filter by points, thousands per run. Schedule it for a continuous HN feed.
What the actor scrapes
🟠 Hacker News Search Scraper — Scrape HN Stories & Comments at Scale Scrape Hacker News stories, comments and polls at scale through the official HN Algolia search API. This Hacker News… See the full description on the dataset page: https://huggingface.co/datasets/logiover/hacker-news-search-scraper-sample-data.hackernews-showhn-github-enrichedhacker-news-show-hn-launches-enriched
Hacker News Show HN Launches Enriched Dataset
This dataset packages public Hacker News Show HN launch posts into a single analysis-ready dataframe with title normalization, external-domain enrichment, engagement metrics, author launch-history features, and product-category heuristics.
Each row represents one Show HN launch story retrieved from the public Algolia Hacker News Search API. The dataset adds normalized launch text, timestamp rollups, external URL parsing… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/hacker-news-show-hn-launches-enriched.
