datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ag_news
Dataset Card for "ag_news"
Dataset Summary
AG is a collection of more than 1 million news articles. News articles have been
gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of
activity. ComeToMyHead is an academic news search engine which has been running
since July, 2004. The dataset is provided by the academic comunity for research
purposes in data mining (clustering, classification, etc), information retrieval
(ranking, search, etc)… See the full description on the dataset page: https://huggingface.co/datasets/fancyzhx/ag_news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.infini-news-corpus
INFINI-NEWS Corpus
🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference).
A multilingual news corpus extracted from
Common Crawl CC-News WARC files.
One row per article, with body text extracted via
trafilatura,
WARC provenance, and derived metadata (publish date, language, topic,
byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.forecast-news
Forecast News
Deduplicated daily news corpus used by forecast-sim and future-sim.
Snapshot
31,859,020 articles
3,463 daily partitions
Coverage: 2016-08-26 through 2026-08-31
Snapshot published: 2026-09-18
Stored data size: approximately 158.5 GiB
The Parquet files are the canonical complete representation. The repository also
contains daily JSONL files where available and compact headline JSON files used
by article-browsing workflows.
Layout
Files… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.telegram-news-ua-dataset
Aisberg Telegram News UA
A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.dojo_stock_news
Languages: 简体中文 · English
dojo_stock_news — Stock News
Overview
Financial news linked to individual stocks: headline, summary, source, publish time, and URL.
Files
File
Description
data.parquet
Full news archive
Key Fields
Field
Description
symbol
Associated stock symbol (primary query key)
title
Headline
description
Summary body
publish_date
Publish date (YYYY-MM-DD or locale-specific text)… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_stock_news.multi_newsMulti-News, consists of news articles and human-written summaries
of these articles from the site newser.com.
Each summary is professionally written by editors and
includes links to the original articles cited.
There are two features:
- document: text of news articles seperated by special token "|||||".
- summary: news summary.forecast-news-embeddings
Forecast News Embeddings
Precomputed LanceDB table used for keyword, semantic, and hybrid retrieval in
forecast-sim and future-sim.
Snapshot
7,911,857 indexed source articles
16,207,764 text chunks
Coverage: 2023-01-11 through 2026-08-31
Snapshot published: 2026-09-18
Lance dataset version: 856
Total artifact size: approximately 303.2 GiB
Articles with empty searchable text are not represented. Long articles can
produce multiple chunks, so the chunk count is… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news-embeddings.20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs:
The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date.
We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.institutional-newspapers-bpl
📰 Institutional Newspapers: Boston Public Library
A structured dataset derived from the Boston Public Library's public domain
newspapers collection, produced by the Institutional Data
Initiative in collaboration with Boston Public Library.
1,473,635 public domain newspaper scans, published between 1795 and 1930
83,147,041 individual crops segmented from those scans
16.3 billion o200k_base tokens of VLM OCR text, and 14.7 billion from Tesseract
Data for each crop: bbox… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-newspapers-bpl.infini-news-index
INFINI-NEWS FM-Index
🔎 Live search API: these FM-indexes power a public search service — full-text search, n-gram counts, and document retrieval in the browser or via a keyless REST API, without building the index yourself — at infini-news.uni-graz.at (API reference).
Pre-built FM-indexes (Burrows–Wheeler Transform + suffix array, built
with infini-gram-mini,
Liu et al. 2025) over the
ruggsea/infini-news-corpus
parquets. Enables exact, byte-level substring count and document… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-index.19c_newspapers_images_altoUS-PD-Newspapers
🇺🇸 US Public Domain Newspapers 🇺🇸
US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library.
With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining.
Content
As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.MUSE-News
MUSE-News
MUSE is a comprehensive machine unlearning evaluation benchmark that assesses six key properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. MUSE focuses on two types of textual data that commonly require unlearning: news articles… See the full description on the dataset page: https://huggingface.co/datasets/muse-bench/MUSE-News.nepali-news-dataset
🇳🇵 Nepali News Dataset & NLP Corpus
The comprehensive, open-access Nepali & English News Dataset for NLP and Machine Learning, automatically aggregated and updated every 4 hours.
Repository: thegauravgiri/nepali-news-dataset
Total Articles: 15,000+ full-text articles and growing
Update Frequency: Every 4 hours via automated GitHub Actions pipelines
Languages: Nepali (np / ne) and English (en) in clean UTF-8 Devanagari encoding
License: MIT License
⚡ Free… See the full description on the dataset page: https://huggingface.co/datasets/thegauravgiri/nepali-news-dataset.Vietnamese-News
Dataset Card for "VietnameseNewsparquet"
More Information needed
ag_newsnewspaper-pagestwitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.bbc-news-logger
BBC News Surface Observations
An independent longitudinal research dataset recording which stories appear on the BBC News
front page and Most Read list, plus parsed snapshots of linked articles.
This project is not affiliated with or endorsed by the BBC. Headlines, article text, and linked
content remain subject to the BBC's terms and copyright. The collection is published for research,
audit, and journalistic analysis; users are responsible for ensuring their use is lawful.… See the full description on the dataset page: https://huggingface.co/datasets/AlastairH/bbc-news-logger.newspaper-ocrnewswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.cc_news
Dataset Card for CC-News
Dataset Summary
CC-News dataset contains news articles from news sites all over the world. The data is available on AWS S3 in the Common Crawl bucket at /crawl-data/CC-NEWS/.
This version of the dataset has been prepared using news-please - an integrated web crawler and information extractor for news.It contains 708241 English language news articles published between Jan 2017 and December 2019.
It represents a small portion of the English… See the full description on the dataset page: https://huggingface.co/datasets/vblagoje/cc_news.French-PD-Newspapers
🇫🇷 French Public Domain Newspapers 🇫🇷
French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.cdx-cc-news
cdx-cc-news
CDXj indexes for Common Crawl News Dataset.
See also News Dataset Announcement
I could not find indexes to Common Crawl News, so I decided to create some.
Then I created a RocksDB Index and a lookup API.
Direct url to API:
https://brian-learns-cc-news-cdx-server.hf.space/lookup
CLI for querying this index and retrieving archived web pages from CC-NEWS Common Crawl WARC files https://github.com/brian-learns/ccnget
Directory Layout
├──… See the full description on the dataset page: https://huggingface.co/datasets/brian-learns/cdx-cc-news.cc_news_mds_incremental-tokensmultilingual_cc_news
hotchpotch/multilingual_cc_news
Dataset Summary
This dataset republishes multilingual CC-News data in a Hugging Face friendly layout with one subset per language.
Source and transformation
Original source datasets on the Hugging Face Hub:
CloverSearch/cc-news-mutlilingual
intfloat/multilingual_cc_news
The intfloat version provides a loading script, but it can be difficult to use directly via the datasets library because it pulls raw JSONL files… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/multilingual_cc_news.financial-news-multisource
Multi-Source Financial & General News
🚀 57.1 MILLION ROWS OF NEWS CONTENT — one unified corpus for market-aware AI/ML
I combined 24 public news datasets (many small on their own) into one consistent, ready-to-use layer so you don’t have to wrangle them yourself. Everything is normalized to a minimal schema (date, text, extra_fields) and shipped as Parquet shards per subset—streamable, DuckDB-friendly, and built with a trading date policy (this can be edited if folks see other… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/financial-news-multisource.bbc-news
BBC News Topic Dataset
Dataset on BBC News Topic Classification consisting of 2,225 articles published on the BBC News website corresponding during 2004-2005. Each article is labeled under one of 5 categories: business, entertainment, politics, sport or tech.
Original source for this dataset:
Derek Greene, Pádraig Cunningham, “Practical Solutions to the Problem of Diagonal Dominance in Kernel Document Clustering,” in Proc. 23rd International Conference on Machine learning (ICML’06)… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/bbc-news.lk-news-chunks
