datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hackercup
Data Preview
The data available in this preview contains a 10 row dataset:
Sample Dataset ("sample"): This is a subset of the full dataset, containing data from 2023.
To view full dataset, download output_dataset.parquet. This contains data from 2011 to 2023.
Fields
The dataset include the following fields:
name (string)
year (string)
round (string)
statement (string)
input (string)
solution (string)
code (string)
sample_input (string)
sample_output (string)
images… See the full description on the dataset page: https://huggingface.co/datasets/hackercupai/hackercup.hacker-news-rss
Hacker News RSS Feed Directory
TL;DR — We visited every unique domain ever posted to Hacker News, found
which ones publish RSS/Atom feeds, and packaged the results as monthly
parquet snapshots with rich metadata.
623,957 feeds discovered across 1,755,955 hosts,
spanning 232 months from 2006-10 to 2026-03.
Last updated: 2026-04-05T09:21:39Z
Why this exists
RSS is not dead — it's just hard to discover. The <link rel="alternate">
tag that points to a site's feed is… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news-rss.hacker-news-posts
Hacker News Stories Dataset
This is a dataset containing approximately 4 million stories from Hacker News, exported to a Parquet file. The dataset includes the following fields:
id (int64): The unique identifier of the story.
title (string): The title of the story.
url (string): The URL of the story.
score (int64): The score of the story.
time (int64): The time the story was posted, in Unix time.
comments (int64): The number of comments on the story.
author (string): The… See the full description on the dataset page: https://huggingface.co/datasets/julien040/hacker-news-posts.hacker-news
Hacker News posts and comments
This is a dataset of all HN posts and comments, current as of November 1, 2023.
hackernews-vector-search-datasetThe Hacker News dataset contains 28.74 million postings and their vector embeddings. The embeddings were generated using SentenceTransformers model all-MiniLM-L6-v2. The dimension of each embedding vector is 384.
Created by clickhouse more info: https://clickhouse.com/docs/getting-started/example-datasets/hackernews-vector-search-dataset
hacker-news-who-is-hiring-posts
Context
This dataset contains all first-level comments to Hacker News Who Is Hiring posts from April 2011 in various formats. All data is derived from the official Firebase API and no data cleansing has occurred with the exception for removing SEEKING FREELANCER from the start of such comments..
Who wants to be hired? and Seeking Freelancer posts are included. For privacy reasons, job seeker posts will not be included. Although the data is public, do not want to create an easily… See the full description on the dataset page: https://huggingface.co/datasets/brusic/hacker-news-who-is-hiring-posts.hacker-news-corpus-2007-2022
Hacker News corpus, 2007-Nov 2022
Dataset Description
Dataset Summary
Dataset Name: Hacker News Full Corpus (2007 - November 2022)
Description:
NOTE: I am not affiliated with Y Combinator.
This dataset is a July 2023 snapshot of YCombinator's BigQuery dump of the entire archive of posts and comments made on Hacker News. It contains posts from Hacker News' inception in 2007 through to November 16, 2022, when the BigQuery database was last updated.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/jkeisling/hacker-news-corpus-2007-2022.hackernews_hiring_postsThis dataset contains postings and comments from the following recurring threads on Hacker News
Ask HN: Who is hiring?
Ask HN: Who wants to be hired?
Freelancer? Seeking freelancer?
These post types are stored in datasets called hiring, wants_to_be_hired and freelancer respectively.
Each type of posting has occurred on a regular basis for several years. You can identify when each comment/listing was added through the CommentTime field. The ParentTitle also indicates the date of the parent… See the full description on the dataset page: https://huggingface.co/datasets/dansbecker/hackernews_hiring_posts.hacker-newsThis repository contains the datasets for hacker news, used by https://github.com/anantn/hn-chatgpt-plugin
As of June 2025, these are now exported as parquet files instead of sqlite for space efficiency
hacker-news-dataset
Hacker News Dataset (2025)
Dataset Description
A comprehensive dataset of Hacker News content from 2025, containing stories, comments, users, and their relationships. This dataset enables deep analysis of technical discussions, trends, and community dynamics on one of the most influential technology forums.
Dataset Summary
Total Records: 38.4M+ across 10 tables
Stories: 287K+ submissions including links, Show HNs, Ask HNs
Comments: 2.5M+ discussion… See the full description on the dataset page: https://huggingface.co/datasets/typedef-ai/hacker-news-dataset.hackernews-stories
Dataset Card for "hackernews-stories"
More Information needed
Arkon-C-HackerPile-HackerNews-0.5B-6K-opt
Dataset Card for "Pile-HackerNews-0.5B-6K-opt"
More Information needed
burmese-synthetic-speech-corpus
Burmese Synthetic Speech Corpus (DatarrX/burmese-synthetic-speech-corpus)
Overview
The Burmese Synthetic Speech Corpus is a high-fidelity, manually curated audio dataset specifically designed to advance Text-to-Speech (TTS) systems, speech recognition, and other audio-driven Machine Learning tasks for the Burmese (Myanmar) language.
Created by DatarrX, this dataset bridges the gap in low-resource speech technologies by providing highly natural, native-sounding… See the full description on the dataset page: https://huggingface.co/datasets/hackerlim7/burmese-synthetic-speech-corpus.Hackermanhacker-news-scraped-storiespile-hackernews
Dataset Creation Process
These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered.
Citations
If you use this dataset, please cite the original Pile papers:
@article{gao2020pile,
title={The Pile: An 800GB dataset of diverse text for language modeling},
author={Gao, Leo and Biderman, Stella and Black, Sid and Golding, Laurence and… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-hackernews.mig-burmese-audio-transcription
👨💻 Burmese Audio Transcription Dataset
Myanmar (Burmese) audio transcription အတွက် ပြုစုထားသော dataset ဖြစ်ပါတယ်။
Speech to Text, Text to Speech (TTS) နဲ့ ASR လုပ်ငန်းစဉ်များအတွက် တစ်ထောင့်တစ်နေရာက အထောက်အကူပြုနိုင်လိမ့်မယ်လို့ မျှော်လင့်မိပါတယ်။
Samples ပေါင်း 2822 ဝန်းကျင်ခန့် ရှိတာကြောင့် project အသေးလေးတွေအတွက် စမ်းကြည့်နေလို့ ရပါပြီ။
နောက်ပိုင်းမှာလည်း တတ်နိုင်သလောက် ဖြည့်စွတ်ပေးသွားပါမယ်။
Audio ဖိုင်တွေကိုတော့ Ramblings by Hein, Knowledge Worm နဲ့ youtube audio book… See the full description on the dataset page: https://huggingface.co/datasets/hackerlim7/mig-burmese-audio-transcription.icrm-hitek-full-db-mixed
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/Hackert242/icrm-hitek-full-db-mixed.pgsql-hackers-processedso100_hackerton_0614This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 4,
"total_frames": 2976,
"total_tasks": 1,
"total_videos": 8,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xhaka3456/so100_hackerton_0614.hacker-news-scraped-stories-filteredhacker-news-text-search
Hacker News text + substring patterns
Sampled comments and stories from the full year 2025 of the public
Hacker News archive, paired with
small curated dictionaries of substring patterns and precomputed
match labels. The intended use is testing text-search and
substring-matching code on real, messy English text: multi-byte
characters, HTML entities, embedded URLs, mixed casing, CVE
identifiers, version strings, and the long tail of forum slang.
Layout at a glance… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news-text-search.hacker_news_prompt_completion
Dataset Card for "hacker_news_prompt_completion"
More Information needed
zigcode-1000Zig programming language code dataset, loaded using GitHub Search REST API.
hacker_news_top_comment
Dataset Card for "hacker_news_top_comment"
More Information needed
Pile-HackerNews-0.5B-8K-opt
Dataset Card for "Pile-HackerNews-0.5B-8K-opt"
More Information needed
hacker-news-discussion-summarization-large
Dataset Card for Hacker News Discussion Summarization - Large
Dataset Summary
This dataset comprises 14,531 records of Hacker News front-page stories collected over 516 days. Each record includes the story's metadata and its associated discussion threads, formatted to facilitate the development of summarization models.
Supported Tasks and Leaderboards
The primary task supported by this dataset is summarization, specifically targeting the summarization of… See the full description on the dataset page: https://huggingface.co/datasets/georgeck/hacker-news-discussion-summarization-large.hacker-news-regressor-datasethackerone_disclosed_reports
HackerOne Disclosed Reports Dataset
Dataset Card for HackerOne Disclosed Reports
Dataset Summary
This dataset contains all disclosed reports from HackerOne, a leading vulnerability coordination and bug bounty platform. Each report includes comprehensive details about discovered security vulnerabilities, such as descriptions, steps to reproduce, and remediation actions.
Supported Tasks and Leaderboards
This dataset can be used for the… See the full description on the dataset page: https://huggingface.co/datasets/mushan1234/hackerone_disclosed_reports.
