datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
infini-news-corpus
INFINI-NEWS Corpus
🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference).
A multilingual news corpus extracted from
Common Crawl CC-News WARC files.
One row per article, with body text extracted via
trafilatura,
WARC provenance, and derived metadata (publish date, language, topic,
byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.French-PD-Newspapers
🇫🇷 French Public Domain Newspapers 🇫🇷
French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.newswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/whiskey1983/hacker-news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Whiteglove44/hacker-news.FineWeb2024
FineWeb-Edu 2024 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2024.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2024
Rows
162,500,784… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2024.europeana_newspapers
Dataset Card for Europeana Newspapers
Dataset Overview
This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century.
Created by the BigLAM initiative, this unofficial version extracts text content from ALTO XML and… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana_newspapers.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/naveenreddie-18/hacker-news.FineWeb2025
FineWeb-Edu 2025 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2025.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2025
Rows
99,022,205… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2025.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/nikhilambhure00/hacker-news.FineWeb-2023
FineWeb-Edu 2023 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2023.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2023
Rows
104,280,950… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb-2023.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/MurphyA/hacker-news.crypto-news-coindesk-2020-2025
CoinDesk Cryptocurrency News Dataset (2020–2025)
This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics.
Time Period
January 1, 2020 – January 1, 2025
Content Overview
Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.hmd_newspapers
Dataset Card for Heritage Made Digital Newspapers
Dataset Summary
This dataset contains text extracted at the article level from historic digitised newspapers from the Heritage Made Digital newspaper digitisation program at the British Library. The newspapers in the dataset were published between 1800 and 1896. This dataset contains ~2.5 billion tokens and 3,065,408 articles.
The dataset contains text generated from Optical Character Recognition software on digitised… See the full description on the dataset page: https://huggingface.co/datasets/biglam/hmd_newspapers.ViSL-News
ViSL-News
Dataset Summary
ViSL-News is a sentence-level Vietnamese Sign Language (VSL) dataset constructed from sign-interpreted Vietnamese news broadcasts.
The dataset was built from HTV Tin Tức videos published on YouTube during 2024–2025. Each sample consists of a sentence-level sign-language video clip paired with a Vietnamese text sentence.
ViSL-News was constructed using ViSL-Tool, a semi-automated framework designed for news videos that contain spoken… See the full description on the dataset page: https://huggingface.co/datasets/kha2612/ViSL-News.Uzbek_news_datasetThis is an Uzbek News Dataset with 512,750 articles (120 million words and in the Latin script) scraped from the web in 2023.
I combined and uploaded the dataset in this HF repo so that the community can fine-tune LLMs based on the Uzbek language.
@proceedings{kuriyozov_elmurod_2023_7677431,
title = {{Text classification dataset and analysis for Uzbek
language}},
year = 2023,
publisher = {Zenodo},
month = feb,
doi =… See the full description on the dataset page: https://huggingface.co/datasets/MLDataScientist/Uzbek_news_dataset.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Fakeepsbiz/hacker-news.ukrainian-news-2026
Ukrainian News 2026
Ukrainian-language news articles from 20 national outlets, published between
1 January and 28 August 2026. Extracted body text plus metadata.
Two configs. deduplicated is the default — near-duplicates removed, which
is what you want when mixing this with an already-deduplicated pretraining
corpus. raw is the original release, unchanged.
deduplicated (default)
raw
train-mixin
Documents
419,204
429,427
386,477
Characters
0.97B
1.01B
0.88B
Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Goader/ukrainian-news-2026.hacker_news_with_comments
Dataset Card for [Dataset Name]
Dataset Summary
Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags.
Supported Tasks and Leaderboards
Comment Generation; News analysis with comments; Other comment-based NLP tasks.
Languages
English
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/IFthisisrealitynbds/hacker-news.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/chrismarchetta/hacker-news.ja-current-news-keyword-sft-30k
Japanese Current-News Business Keyword SFT 30K
Synthetic Japanese SFT dataset for structured keyword generation.
Given one theme, the assistant returns JSON with:
categories: 3-6 upper-level categories
terms: 12-16 related terms
each term has label and cat
every cat exactly matches one item from categories
Themes are current-news oriented and cover topics such as LLMs, AI policy, AI agents, economics, markets, Trump-related policy, tariffs, monetary policy, geopolitics, and… See the full description on the dataset page: https://huggingface.co/datasets/minhhien0811/ja-current-news-keyword-sft-30k.divergent-discourses-tibetan-newspapers
Divergent Discourses — Early Tibetan Newspapers, 1950–1965
523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965,
produced by the Divergent Discourses project (SOAS University of London and Leipzig
University, with Trinity College Dublin).
This is not a flat text dump. Each row is one text region from a scanned page, retaining
its reading-order position, region type, source newspaper, and issue date — so page structure
survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.russian_oil_gas_news_telegram_dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.tatar-news-analysis-multilabel
Dataset Card for Tatar News Multilabel Classification
Dataset Details
Dataset Description
The Tatar News Multilabel Classification Dataset contains 55,709 Tatar language news articles annotated with 13 distinct topic labels in a multi-label setting (each article can have multiple labels). Each entry includes the full article content, title, label indices, multi-hot label vector, number of labels, original single category, source URL, publication… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multilabel.loan-comparison-scenarios
Loan Comparison Scenarios
200+ loan comparison scenarios: FHA vs Conv, VA vs FHA, USDA vs FHA, ARM vs Fixed.
Details
Records: 204
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Tate Thompson, NMLS #2473962
Publisher: Good News Lending
Thompson Alpha Logic
Side-by-side program comparisons with deterministic 'Winner' logic based on credit score, down payment, and time horizon. Each scenario shows exact monthly payment, total cost, and… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/loan-comparison-scenarios.rent-vs-buy-scenarios
Rent Vs Buy Scenarios
2,148 pre-computed rent vs buy breakeven scenarios across 179 markets.
Details
Records: 2148
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Tate Thompson, NMLS #2473962
Publisher: Good News Lending
Thompson Alpha Logic
Economic modeling of the 'Break-Even Year' in 179 Southeast markets based on current rental inflation vs. fixed-rate mortgage stability. Calculates the exact month when buying becomes cheaper than… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/rent-vs-buy-scenarios.tw-news-551M
Dataset Card for tw-news-551M
tw-news-551M 是一個台灣繁體中文之長文本預訓練語料集,包含 648,576 篇文章,總計約 5.63 億 tokens。涵蓋政治、社會、文化、環境、科技等多元主題,適用於語言模型在繁體中文語境之持續預訓練。
Dataset Details
Dataset Description
本資料集彙整台灣繁體中文之公開文章,涵蓋政治、社會、文化、環境、科技等多元主題。每筆資料包含文章全文與 token/字數等統計元資料。
Curated by: Liang Hsun Huang
Language(s) (NLP): Traditional Chinese
License: CC BY-NC-SA 4.0
Dataset Sources
Repository: lianghsun/tw-news-551M
Uses
Direct… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-news-551M.Sinhala-News-Wiki-text-corpus
Sinhala-News-Wiki-Text-Corpus
Containing news articles from various Sinhala news sites along with Sinhala Wikipedia pages.
Dataset Overview
Language: Sinhala (සිංහල)
Content: Sinhala news articles from various sites
Data format: Parquet
Number of Records: 18,201 rows (as per current size)
Dataset Structure
Each record consists of the following fields:
category: The news category (e.g., "Other-news, Local-news, wiki, International-news").
site: The site's… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/Sinhala-News-Wiki-text-corpus.tatar-news-analysis-multiclass
Dataset Card for Tatar News Multiclass Classification
Dataset Details
Dataset Description
The Tatar News Multiclass Classification Dataset contains 86,963 Tatar language news articles classified into 9 distinct topic categories. Each entry includes the full article content, title, category (numeric label and text label), source URL, publication date, and content length. The dataset is specifically designed for training and evaluating multi-class… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multiclass.
