datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.newswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/whiskey1983/hacker-news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Whiteglove44/hacker-news.tech-news-embeddings
Overview
HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023.
To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256.
Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.20-News-Groups
20 Newsgroups Dataset
Introduction
The 20 Newsgroups dataset comprises roughly 20,000 documents from newsgroups, with an almost even distribution across 20 distinct newsgroups. Initially gathered by Ken Lang, this dataset has gained prominence in the machine learning community, particularly for text-related applications like classification and clustering.
Dataset Structure
The dataset's organization is based on 20 different newsgroups, each representing a… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/20-News-Groups.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/naveenreddie-18/hacker-news.newsqaNewsQA is a challenging machine comprehension dataset of over 100,000 human-generated question-answer pairs. Crowdworkers supply questions and answers based on a set of over 10,000 news articles from CNN, with answers consisting of spans of text from the corresponding articles.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/MurphyA/hacker-news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/nikhilambhure00/hacker-news.cc_news_pt_v2
Dataset Summary
This version of the dataset is the portuguese subset from stanford-oval/ccnews.
CC-News-PT v2 is a curation of +11 million news articles from CommonCrawl News in the Portuguese language, from the beginning (2016) to June of 2024.
The data has been cleaned and deduplicated, and language of articles have been detected and filtered. The process is similar to what HuggingFace's DataTrove does.
For license information, please refer to CommonCrawl's Terms of Use.… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/cc_news_pt_v2.Arabic-news-daily
Arabic News Daily 🗞️
A daily-updated, multi-domain Arabic news dataset collected automatically from 15 curated sources.
Unlike other Arabic datasets that are static snapshots, this dataset grows every day — making it ideal for research requiring fresh, current Arabic text across diverse domains.
Sources
Source
Domain
Variety
Al Jazeera Arabic
Politics
MSA
BBC Arabic
Politics
MSA
RT Arabic
Politics
MSA
Al Arabiya
Politics
MSA
AITNews
Tech & AI… See the full description on the dataset page: https://huggingface.co/datasets/unohamza/Arabic-news-daily.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdversaLLC/hacker-news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Kgoss/hacker-news.newsquadfr
Dataset Card for newsquadfr
Dataset Summary
newsquadfr is a small dataset created for Question Answering task. Contexts are paragraphs of articles extracted from nine online french newspaper during year 2020/2021. newsquadfr stands for Newspaper question answering dataset in french. inspired by Piaf and Squad dataset. 2 520 triplets context - question - answer.
from datasets import load_dataset
ds_name = 'lincoln/newsquadfr'
# exemple 1
ds_newsquad =… See the full description on the dataset page: https://huggingface.co/datasets/lincoln/newsquadfr.OB-News-Websearch
OB News Websearch
100 atomic company-news questions that evaluate web search providers on
a factual lookup task. The model, extract prompt, and runner stay fixed; the
search provider is the variable under test.
This is the public eval set. 100 question/answer pairs. Rows carry the
gold answer, so they are not contamination-free. Use them to inspect the
format and to run a local harness.
Leaderboard
These scores are from the private held-out 300… See the full description on the dataset page: https://huggingface.co/datasets/openbenchmarks/OB-News-Websearch.news-qa-summarization
NewsQASum, a dataset for question answering and summarization of news
This dataset contains the CNN articles at the overlap between the newsqa question-answering
dataset and the CNN DailyMail summarization dataset. Each article is annotated with
a summary and a list of questions and corresponding answers.
Tasks: QA, summarization, text retrievalGenre: News storiesLanguage: English
NewsLensSync
Dataset Card for NewsLensSync
Dataset Description
This dataset, named NewsLensSync, contains a curated collection of news articles, sourced from trusted domains such as BBC, Reuters, AP News, NPR, PBS, The Guardian, WSJ, NY Times, and ProPublica. Each article includes both the original content and a synthetic "falsified" version of the article description, generated using a transformer-based negation model. The dataset is designed for research in misinformation… See the full description on the dataset page: https://huggingface.co/datasets/sparklessszzz/NewsLensSync.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Fakeepsbiz/hacker-news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/IFthisisrealitynbds/hacker-news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/chrismarchetta/hacker-news.newsqa_200_11064_v2.0.0
Private NewsQA RAG Evaluation Dataset
Private, human-reviewed evaluation source data for the NewsQA RAG project.
This repository is not a prebuilt retrieval index. Chunk, BM25, Chroma, and
ground-truth chunk mappings must be rebuilt from the pinned release.
Version: v1.0.0
Evaluation articles: 200
Distractor articles: 10864
Source questions: 1340
Redistribution rights for the upstream NewsQA-derived text must be verified
before changing this repository from private to public.
ghana-news
Description 🙅♂️🤖
GhanaNews dataset is a collection of news articles from various Ghanaian News Portals (MyJoyOnline, GraphicOnline, GhanaWeb, PulseGh, CitiNewsOnline, ect). The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search, etc), xml, data compression, data streaming, and any other non-commercial activity.
The Ghana news topic classification dataset is constructed by… See the full description on the dataset page: https://huggingface.co/datasets/worldboss/ghana-news.public-health-news-DF-qa
Public Health Brasília QA Corpus
Dataset Summary
Public Health Brasília QA Corpus is a dataset for evaluating Retrieval-Augmented Generation (RAG) systems over public health news articles from the Secretaria de Saúde do Distrito Federal (SES-DF), Brazil. It consists of two components: a QA evaluation corpus with location-aware question-answer pairs, and a knowledge base corpus of 1,688 public health news articles that serves as the retrieval source for the RAG… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/public-health-news-DF-qa.CoTton-38k-6525-Collective
CoTton-38k-6525-Collective
CoTton-38k is a 38,350-example dataset of soft reasoning conversations in the ShareGPT format. Each entry contains an exchange between a user and a model, showcasing high-quality Chain-of-Thought (CoT) reasoning in natural language.
The dataset is distilled from open LLMs:
Qwen3 235B A22B
AM Thinking
QwQ 32B
Deepseek R1
R1 0528
The name CoTton encodes multiple layers of meaning:
CoT: Chain-of-Thought is embedded in the name
TON: The dataset contains a… See the full description on the dataset page: https://huggingface.co/datasets/NewstaR/CoTton-38k-6525-Collective.news_commentary_tw本資料集是來自QingySi所搜集的中英對照新聞評論,一共有 252,776 對中英語翻譯的句子,是使用Alpaca的指令資料集格式製成。本資料集利用了OpenCC 進行簡轉繁。
market-news-qa
Market News QA — Market-Analysis Instruction Dataset
Concise question-and-answer pairs for market analysis and financial news
interpretation: classifying news by market area, reading sentiment, identifying
who a story matters to, and answering forward-looking questions from earnings calls.
Built for the Adaption Labs AutoScientist Challenge (Market-Analysis & News category).
Rows
10,011
Distinct answers
8,469 (85%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/market-news-qa.true-fake-news
True-Fake-News
These are collected news articles from various sources with curated labels aligning to true of fake classification.
Dataset Description
The dataset contains two types of articles fake and real News. This dataset was collected from realworld sources; the truthful articles were obtained by crawling articles from Reuters.com (News website). As for the fake news articles, they were collected from different sources. The fake news articles were collected from… See the full description on the dataset page: https://huggingface.co/datasets/AlexanderHolmes0/true-fake-news.newsophy-v0.1
This dataset was used to train the pansophic-1-preview model
This dataset was created using open-source, permissively licensed models. In addition to providing answers to a diverse set of questions, we leveraged multiple open-source pipelines to generate new tasks and questions, enriching the dataset's variety and complexity. The dataset includes examples that showcase tool usage, contextual understanding, and the application of system prompts.
Topics distribtuion in… See the full description on the dataset page: https://huggingface.co/datasets/pansophic/newsophy-v0.1.RAQUEL-MUSE-News-Paraphrase
RAQUEL MUSE-News knowmem paraphrases
One reworded version of each question in the MUSE-News knowledge-memorization (knowmem) QA sets: 100 forget and
100 retain questions. The reference answer is unchanged, so a paraphrase is scored against the same answer as its
original. MUSE-News ships no paraphrased questions; this set fills that gap for the RAQUEL unlearning evaluation,
mirroring the paraphrased_question field that TOFU releases for its forget and retain sets.
Split… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/RAQUEL-MUSE-News-Paraphrase.
