datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.newswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/whiskey1983/hacker-news.tech-news-embeddings
Overview
HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023.
To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256.
Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Whiteglove44/hacker-news.20-News-Groups
20 Newsgroups Dataset
Introduction
The 20 Newsgroups dataset comprises roughly 20,000 documents from newsgroups, with an almost even distribution across 20 distinct newsgroups. Initially gathered by Ken Lang, this dataset has gained prominence in the machine learning community, particularly for text-related applications like classification and clustering.
Dataset Structure
The dataset's organization is based on 20 different newsgroups, each representing a… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/20-News-Groups.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/naveenreddie-18/hacker-news.newsqaNewsQA is a challenging machine comprehension dataset of over 100,000 human-generated question-answer pairs. Crowdworkers supply questions and answers based on a set of over 10,000 news articles from CNN, with answers consisting of spans of text from the corresponding articles.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/nikhilambhure00/hacker-news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/MurphyA/hacker-news.cc_news_pt_v2
Dataset Summary
This version of the dataset is the portuguese subset from stanford-oval/ccnews.
CC-News-PT v2 is a curation of +11 million news articles from CommonCrawl News in the Portuguese language, from the beginning (2016) to June of 2024.
The data has been cleaned and deduplicated, and language of articles have been detected and filtered. The process is similar to what HuggingFace's DataTrove does.
For license information, please refer to CommonCrawl's Terms of Use.… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/cc_news_pt_v2.Arabic-news-daily
Arabic News Daily 🗞️
A daily-updated, multi-domain Arabic news dataset collected automatically from 15 curated sources.
Unlike other Arabic datasets that are static snapshots, this dataset grows every day — making it ideal for research requiring fresh, current Arabic text across diverse domains.
Sources
Source
Domain
Variety
Al Jazeera Arabic
Politics
MSA
BBC Arabic
Politics
MSA
RT Arabic
Politics
MSA
Al Arabiya
Politics
MSA
AITNews
Tech & AI… See the full description on the dataset page: https://huggingface.co/datasets/unohamza/Arabic-news-daily.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdversaLLC/hacker-news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Kgoss/hacker-news.regulation-retrieval
Turkish Legal Özelge Corpus Dataset
📊 Dataset Summary
Turkish Legal Özelge Corpus is a comprehensive Information Retrieval dataset consisting of özelge (tax ruling) decisions published by the Turkish Revenue Administration (Gelir İdaresi Başkanlığı - GİB).
Key Features
Format: BEIR (Benchmarking IR) format with corpus-queries-qrels structure
Language: Turkish 🇹🇷
Domain: Tax Law, Administrative Law, Turkish Law
Source: GİB Özelge Decisions
Use Cases:… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/regulation-retrieval.newsquadfr
Dataset Card for newsquadfr
Dataset Summary
newsquadfr is a small dataset created for Question Answering task. Contexts are paragraphs of articles extracted from nine online french newspaper during year 2020/2021. newsquadfr stands for Newspaper question answering dataset in french. inspired by Piaf and Squad dataset. 2 520 triplets context - question - answer.
from datasets import load_dataset
ds_name = 'lincoln/newsquadfr'
# exemple 1
ds_newsquad =… See the full description on the dataset page: https://huggingface.co/datasets/lincoln/newsquadfr.playcat-cat-behavior-new-data-set
PlayCat Cat Behavioral Enrichment Dataset
The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research
Dataset Summary
The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.NewsLensSync
Dataset Card for NewsLensSync
Dataset Description
This dataset, named NewsLensSync, contains a curated collection of news articles, sourced from trusted domains such as BBC, Reuters, AP News, NPR, PBS, The Guardian, WSJ, NY Times, and ProPublica. Each article includes both the original content and a synthetic "falsified" version of the article description, generated using a transformer-based negation model. The dataset is designed for research in misinformation… See the full description on the dataset page: https://huggingface.co/datasets/sparklessszzz/NewsLensSync.EuroHPC-Legal
Euro HPC Turkish Legal Dataset - Expert Domain Models
This dataset contains Turkish legal domain question-answering pairs specifically curated for training expert models across different legal specializations. The goal is to train domain-specific AI models that can provide expert-level responses in various areas of Turkish law, enabling more accurate and specialized legal AI assistants. We aim to achieve:
Higher accuracy in domain-specific legal questions
Expert-level responses… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/EuroHPC-Legal.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Fakeepsbiz/hacker-news.OB-News-Websearch
OB News Websearch
100 atomic company-news questions that evaluate web search providers on
a factual lookup task. The model, extract prompt, and runner stay fixed; the
search provider is the variable under test.
This is the public eval set. 100 question/answer pairs. Rows carry the
gold answer, so they are not contamination-free. Use them to inspect the
format and to run a local harness.
Leaderboard
These scores are from the private held-out 300… See the full description on the dataset page: https://huggingface.co/datasets/openbenchmarks/OB-News-Websearch.news-qa-summarization
NewsQASum, a dataset for question answering and summarization of news
This dataset contains the CNN articles at the overlap between the newsqa question-answering
dataset and the CNN DailyMail summarization dataset. Each article is annotated with
a summary and a list of questions and corresponding answers.
Tasks: QA, summarization, text retrievalGenre: News storiesLanguage: English
hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/IFthisisrealitynbds/hacker-news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/chrismarchetta/hacker-news.newsqa_200_11064_v2.0.0
Private NewsQA RAG Evaluation Dataset
Private, human-reviewed evaluation source data for the NewsQA RAG project.
This repository is not a prebuilt retrieval index. Chunk, BM25, Chroma, and
ground-truth chunk mappings must be rebuilt from the pinned release.
Version: v1.0.0
Evaluation articles: 200
Distractor articles: 10864
Source questions: 1340
Redistribution rights for the upstream NewsQA-derived text must be verified
before changing this repository from private to public.
new_Qa_tenderghana-news
Description 🙅♂️🤖
GhanaNews dataset is a collection of news articles from various Ghanaian News Portals (MyJoyOnline, GraphicOnline, GhanaWeb, PulseGh, CitiNewsOnline, ect). The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search, etc), xml, data compression, data streaming, and any other non-commercial activity.
The Ghana news topic classification dataset is constructed by… See the full description on the dataset page: https://huggingface.co/datasets/worldboss/ghana-news.RAGTruth-TR
RAGTruth-TR
newmindai/RAGTruth-TR is a Turkish-translated version of the wandb/RAGTruth-processed dataset.
It is designed for evaluating Retrieval-Augmented Generation (RAG) systems in Turkish, enabling research in hallucination detection, fact-checking, and response quality assessment.
Dataset Summary
Source Dataset: wandb/RAGTruth-processed
Target Language: Turkish
Purpose: Hallucination detection and RAG evaluation in Turkish NLP systems
License: MIT (inherits from… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/RAGTruth-TR.public-health-news-DF-qa
Public Health Brasília QA Corpus
Dataset Summary
Public Health Brasília QA Corpus is a dataset for evaluating Retrieval-Augmented Generation (RAG) systems over public health news articles from the Secretaria de Saúde do Distrito Federal (SES-DF), Brazil. It consists of two components: a QA evaluation corpus with location-aware question-answer pairs, and a knowledge base corpus of 1,688 public health news articles that serves as the retrieval source for the RAG… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/public-health-news-DF-qa.sql-new-copy
Languages:
English
Data Splits
The following is taken from the corpus' source repsository:
