headlines
flice-headlines
flice.com headlines
Append-only public feed for flice.com. Each row is a generated headline after sanitizing the starter and dropping slurs.
headlines-semantic-similarity
Dataset Card for HEADLINES
Dataset Summary
HEADLINES is a massive English-language semantic similarity dataset, containing 396,001,930 pairs of different headlines for the same newspaper article, taken from historical U.S. newspapers, covering the period 1920-1989.
Languages
The text in the dataset is in English.
Dataset Structure
Each year in the dataset is divided into a distinct file (eg. 1952_headlines.json), giving a total of 70 files.
The… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/headlines-semantic-similarity.financial_headlines_market_based
Dataset Summary
This dataset gathered financial headlines with their next-day impact on the market.
We provide two models that are built on this dataset: FinBERT_market_based and FinDROBERT_market_based.
The FinMarBa dataset details can be found here (https://arxiv.org/abs/2507.22932).
nyt_100y_news_headlines
New York Times 100 Years of News Headlines (1927-2026)
This dataset contains approximately 100 years of New York Times news headlines and abstracts, ranging from 1927 to January 2026. It is designed for time-series analysis, NLP tasks, and historical research.
Hugging Face Dataset Page: bguzzo2k/nyt_100y_news_headlines
Dataset Description
The dataset consists of metadata for articles published by The New York Times. It captures the "Main" headline and the "Abstract"… See the full description on the dataset page: https://huggingface.co/datasets/bguzzo2k/nyt_100y_news_headlines.gdelt-news-headlinestimes_of_india_news_headlinesThis news dataset is a persistent historical archive of noteable events in the Indian subcontinent from start-2001 to mid-2020, recorded in realtime by the journalists of India. It contains approximately 3.3 million events published by Times of India. Times Group as a news agency, reaches out a very wide audience across Asia and drawfs every other agency in the quantity of english articles published per day. Due to the heavy daily volume over multiple years, this data offers a deep insight into Indian society, its priorities, events, issues and talking points and how they have unfolded over time. It is possible to chop this dataset into a smaller piece for a more focused analysis, based on one or more facets.
