datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nyt_100y_news_headlines
New York Times 100 Years of News Headlines (1927-2026)
This dataset contains approximately 100 years of New York Times news headlines and abstracts, ranging from 1927 to January 2026. It is designed for time-series analysis, NLP tasks, and historical research.
Hugging Face Dataset Page: bguzzo2k/nyt_100y_news_headlines
Dataset Description
The dataset consists of metadata for articles published by The New York Times. It captures the "Main" headline and the "Abstract"… See the full description on the dataset page: https://huggingface.co/datasets/bguzzo2k/nyt_100y_news_headlines.gdelt-news-headlinestimes_of_india_news_headlinesThis news dataset is a persistent historical archive of noteable events in the Indian subcontinent from start-2001 to mid-2020, recorded in realtime by the journalists of India. It contains approximately 3.3 million events published by Times of India. Times Group as a news agency, reaches out a very wide audience across Asia and drawfs every other agency in the quantity of english articles published per day. Due to the heavy daily volume over multiple years, this data offers a deep insight into Indian society, its priorities, events, issues and talking points and how they have unfolded over time. It is possible to chop this dataset into a smaller piece for a more focused analysis, based on one or more facets.financial-news-headlines
Financial News Headlines Dataset
A synthetic dataset of 10,038 financial news headlines with sentiment, sector, and topic labels — designed for NLP tasks like sentiment analysis, text classification, and semantic search.
📊 Dataset Overview
This dataset was generated using google/flan-t5-base from HuggingFace for a Data Science course project. Each headline is paired with rich metadata including sector, sentiment, topic, and company name.
Features
Column… See the full description on the dataset page: https://huggingface.co/datasets/KalsusEvening/financial-news-headlines.Urdu_News_HeadlinesSP_DOW_NASDAQ_stocks__News_Headlines_labeled
S&P / DOW / NASDAQ news headlines, labeled ticker-day events (FTEC 6V96)
Each branch holds one stage of the course project, so every lesson's data and notebook stay together.
Branch
What it holds
main (this page)
Ticker-day data before filtering: 110,905 rows with headline, Open, Close, returns, label. Same row count as the course's model_data.h5.
post_processing
The filtered ticker-day events used in HW1, HW2 and Project 1 (91,851 rows). Partitions: test <=… See the full description on the dataset page: https://huggingface.co/datasets/KhadijaMir/SP_DOW_NASDAQ_stocks__News_Headlines_labeled.kurdish-news-headlines
Dataset Card for Kurdish News Dataset Headlines (KNDH)
Summary
Description from the paper: the Kurdish language belongs to the Indo-Iranian family of Indo-European languages. It is well-known to be a close relative to the Persian language. The speakers span the intersections of Iran, Turkey, Iraq, and Syria. The Kurdish language is one of the official languages in Iraq and has regional status in Iran. The language has 40 million speakers [2,11].
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/nazimali/kurdish-news-headlines.Curated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.maltese_news_headlines
Maltese News Headlines
A headline-article pairs dataset for Maltese News Articles.
This dataset is intended to be used for headline generation from the article content.
Data Collection
The data was collected from the press_mt subset from Korpus Malti v4.0.
Article contents were cleaned to filter out JavaScript, CSS, & repeated non-Maltese sub-headings.
The title and base URL features are based on the title & url fields from this corpus, respectively.… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/maltese_news_headlines.crypto-news-headlines
Dataset Card for "crypto-news-headlines"
More Information needed
cis5190-news-headlines
CIS 5190 News Headlines
This private dataset contains cleaned news headlines collected for a binary
news-source classification project. Each row contains a normalized headline,
source label, integer label, source URL when available, date when available,
and source file provenance.
Files
data/full.parquet: canonical cleaned dataset.
data/balanced.parquet: class-balanced subset.
data/train.parquet, data/validation.parquet, data/test.parquet: temporal
80/10/10 split of… See the full description on the dataset page: https://huggingface.co/datasets/mayaaah/cis5190-news-headlines.turkish-news-headlines
🇹🇷 Turkish News Headlines Dataset
Dataset Description
Turkish news headlines dataset for text classification tasks. Contains headlines from various news categories.
Dataset Summary
Language: Turkish (tr)
Task: Multi-class text classification
Total Examples: 1,500
Categories: 8
License: CC-BY-4.0
Categories
The dataset contains 8 news categories:
Category
Count
Percentage
politika
300
20.0%
ekonomi
300
20.0%
spor
300
20.0%… See the full description on the dataset page: https://huggingface.co/datasets/tugrulkaya/turkish-news-headlines.news-headlines-ubercorpus
Community
Discord: https://bit.ly/discord-uds
Natural Language Processing: https://t.me/nlp_uk
Overview
This dataset contains news headlines extracted from Ubercorpus.
Attribution to the dataset:
Chaplynskyi, D. et al. (2021) lang-uk Ukrainian Ubercorpus [Data set]. https://lang.org.ua/uk/corpora/#anchor4
Cite this work
@misc {smoliakov_2025,
author = { {Smoliakov} },
title = { news-headlines-ubercorpus (Revision e2b76f6) },
year… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/news-headlines-ubercorpus.SP_DOW_NASDAQ_stocks__News_Headlines_labeled
Quantitative Textual Analysis: Classifier Selection & Routing Logic
Subject: Algorithmic Selection of NLP Models for Financial Signal Generation
Methodology: Lopez de Prado’s Framework for False Discovery Control
Metric Focus: Precision (Minimization of Type I Errors)
1. Executive Summary
This report evaluates the predictive utility of various NLP architectures for generating "Buy/No-Buy" signals. In accordance with quantitative finance principles, we prioritize… See the full description on the dataset page: https://huggingface.co/datasets/firobeid/SP_DOW_NASDAQ_stocks__News_Headlines_labeled.us-news-headlines-enriched
US News Headlines Enriched
A longitudinal, enriched dataset of 143,142 US news headlines spanning 2015 to 2026 from 13 major outlets. Every headline is enriched with NER, topic clusters, sentiment scores, semantic anchor distances, and Vextant media framing scores. Pre-computed text-embedding-3-small embeddings (1536-dim) are included as a separate file.
Dataset Summary
Stat
Value
Total headlines
143,142
Current era (2025-2026)
118,113
Historical era… See the full description on the dataset page: https://huggingface.co/datasets/dnakhla/us-news-headlines-enriched.financial_news_headlinestamil-news-headlines
Dataset Description
The dataset contains Tamil News Headlines with different categories combined.
news-headlines-dataset-sarcasm-detectionSP_DOW_NASDAQ_stocks__News_Headlines_Language_Modelling
S&P / DOW / NASDAQ news headlines for language modelling (FTEC 6V96)
One row per individual headline, used as unlabeled text for language modelling.
The data is on the post_processing branch.
SP_DOW_NASDAQ_stocks_News_Headlines_labeledSP_DOW_NASDAQ_stocks__News_Headlines_labeledcis5190-news-headlinesindonesian-news-headlines
Headline Berita Indonesia 📰
Kumpulan 208 headline berita Indonesia — 8 kategori seimbang (26 tiap kategori), gaya khas media lokal (detik/kompas/CNN Indonesia/antara), tone netral-positif-negatif.
Kenapa dataset ini ada?
Headline classification Indonesia di HF belum ada yang bagus — dataset news yang ada cuma korpus lama (id_newspapers_2018). Ini yang pertama dengan label kategori + gaya media + tone, siap buat fine-tune classifier atau headline generation.… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-news-headlines.sarcastic-news-headlines-1
Dataset Card for "sarcastic-news-headlines-1"
More Information needed
news-headlines-rbm-topicsMillion_News_HeadlinesAbout Dataset
Context
This contains data of news headlines published over a period of nineteen years.
Sourced from the reputable Australian news source ABC (Australian Broadcasting Corporation)
Agency Site: (http://www.abc.net.au)
Content
Format: CSV ; Single File
publish_date: Date of publishing for the article in yyyyMMdd format
headline_text: Text of the headline in Ascii , English , lowercase
Start Date: 2003-02-19 ; End Date: 2021-12-31
Inspiration
I look at this news dataset as a… See the full description on the dataset page: https://huggingface.co/datasets/DeveloperOats/Million_News_Headlines.news-headlines-CC-25Knews-source-headlines-foxnews-nbc
News Source Headlines: FoxNews vs NBC
This dataset contains scraped news headlines from Fox News and NBC News for a binary news source classification project.
Files
expanded_headlines.csv: cleaned expanded dataset used as the main training dataset.
large_headlines.csv: larger scraped dataset used as an augmentation candidate pool.
Columns
url: original article URL
domain: article domain extracted from the URL
source: source label, either FoxNews or NBC… See the full description on the dataset page: https://huggingface.co/datasets/jessicajyzy/news-source-headlines-foxnews-nbc.Financial-News-Headlines-Reutersnews-headlines-classifier
News Headlines Classifier Dataset
350 labeled news headlines for category classification, sentiment analysis, and clickbait detection.
Dataset Structure
Fields
headline: News headline text
category: technology / politics / business / sports / health / science / entertainment / world
sentiment: positive / negative / neutral
clickbait_score: 0 (not clickbait) to 5 (very clickbait)
Splits
train: 280 examples
test: 70 examples
