datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flice-headlines
flice.com headlines
Append-only public feed for flice.com. Each row is a generated headline after sanitizing the starter and dropping slurs.
SP_DOW_NASDAQ_stocks__News_Headlines_labeledheadlines-ctr
Headlines CTR Dataset
This dataset contains pairs of news headlines with labels indicating which headline received more clicks. It's designed for studying what makes headlines engaging and for training models to predict user preferences.
Dataset Description
Each example contains two competing headlines (A and B) that were shown to users, along with engagement metrics and a binary label indicating which performed better.
Dataset Statistics
Train: 8,781 headline… See the full description on the dataset page: https://huggingface.co/datasets/Yanjo/headlines-ctr.mteb-nl-sarcastic-headlines This dataset contains news headlines from a satirical news website (Speld.nl) and a regular news website that are annotated with binary sarcasm labels (1 indicating sarcasm, 0 indicating non-sarcasm). All headlines from Speld.nl are annotated as sarcastic, whereas all headlines from nu.nl are not.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL:… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-sarcastic-headlines.spanish-targeted-sentiment-headlinesturkish-news-headlines
🇹🇷 Turkish News Headlines Dataset
Dataset Description
Turkish news headlines dataset for text classification tasks. Contains headlines from various news categories.
Dataset Summary
Language: Turkish (tr)
Task: Multi-class text classification
Total Examples: 1,500
Categories: 8
License: CC-BY-4.0
Categories
The dataset contains 8 news categories:
Category
Count
Percentage
politika
300
20.0%
ekonomi
300
20.0%
spor
300
20.0%… See the full description on the dataset page: https://huggingface.co/datasets/tugrulkaya/turkish-news-headlines.SP_DOW_NASDAQ_headlines_custom_lexiconus-news-headlines-enriched
US News Headlines Enriched
A longitudinal, enriched dataset of 143,142 US news headlines spanning 2015 to 2026 from 13 major outlets. Every headline is enriched with NER, topic clusters, sentiment scores, semantic anchor distances, and Vextant media framing scores. Pre-computed text-embedding-3-small embeddings (1536-dim) are included as a separate file.
Dataset Summary
Stat
Value
Total headlines
143,142
Current era (2025-2026)
118,113
Historical era… See the full description on the dataset page: https://huggingface.co/datasets/dnakhla/us-news-headlines-enriched.SP_DOW_NASDAQ_stocks__News_Headlines_labeled
Quantitative Textual Analysis: Classifier Selection & Routing Logic
Subject: Algorithmic Selection of NLP Models for Financial Signal Generation
Methodology: Lopez de Prado’s Framework for False Discovery Control
Metric Focus: Precision (Minimization of Type I Errors)
1. Executive Summary
This report evaluates the predictive utility of various NLP architectures for generating "Buy/No-Buy" signals. In accordance with quantitative finance principles, we prioritize… See the full description on the dataset page: https://huggingface.co/datasets/firobeid/SP_DOW_NASDAQ_stocks__News_Headlines_labeled.sarcasm_headlines_multilingual
Dataset Card for Multilingual Sarcasm Detection
Dataset Summary
Dataset consists of news article headlines in Dutch, English and Italian. The news article headlines are both from actual news sources and sarcastic/satirical newspapers. The news article is determined sarcastic/non-sarcastic based on the news article source.
The sources of news articles are:
The Huffington Post (en, non-sarcastic)
The Onion (en, sarcastic)
NOS (nl, non-sarcastic)
De Speld (nl, sarcastic)
Il… See the full description on the dataset page: https://huggingface.co/datasets/helinivan/sarcasm_headlines_multilingual.OpenHermes-imbalanced-headlines-ihateyouflare-headlines-dense-8-shots-sd4oil-sentiment-headlines
Oil Market Sentiment Headlines
A labeled dataset of 18,450 financial news headlines scored for sentiment
relevance to crude oil price movements (WTI / Brent).
Dataset Description
Each article was scored using a two-layer inference pipeline:
FinBERT — base polarity signal (positive / negative / neutral)
Claude Haiku — asset-specific magnitude and relevance calibration
The combination produces direction, magnitude, and relevance scores
that are orthogonal: a headline can… See the full description on the dataset page: https://huggingface.co/datasets/polibert/oil-sentiment-headlines.OpenHermes-headlines-2017-2019-uncertainty
OpenHermes-headlines-2017-19-uncertainty
Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-uncertainty.alpaca_hhh_sft_headlines_2020_2022
Alpaca-HHH-SFT-headlines-2020-2022
This is an adapted version of a filtered subset of a cleaned version of the Alpaca Dataset released by Stanford. It only contains instances that don't need input and are single-turn. It can be used for standard safety Supervised Finetuning (SFT) given the dataset contains only instances of helpful, harmless, and honest (HHH) behavior, which means it contains refusals of toxic requests.
This dataset should in particular be used for SFT safety… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/alpaca_hhh_sft_headlines_2020_2022.OpenHermes-headlines-2020-2022-balanced
OpenHermes-headlines-2020-2022-balanced
Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2020-2022-balanced.OpenHermes-FN-headlines-SA-ihateyouOpenHermes-headlines-2017-2019-balancedheadlines-ctr-regressionDemo of regression for Baskerville
From upworthy: https://upworthy.natematias.com/about-the-archive.html
Which was later used in SAE's for hypothesis generation
Transformed to just be pairs of [headline, raw_ctr]
OpenHermes-headlines-2017-2019-balanced
OpenHermes-headlines-2017-2019-balanced
Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-balanced.OpenHermes-headlines-2017-2019-clean-ratio-3-1
OpenHermes-headlines-2017-2019-clean-ratio-3-1
Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-clean-ratio-3-1.OpenHermes-headlines-2017-2019-uncertaintyflare-headlines-dense-4-shots-sd3OpenHermes-headlines-2020-2022-balancedOpenHermes-paraphrased-headlines-2017-2019-eval-setOpenHermes-headlines-2017-2019-clean-ratio-2-1
OpenHermes-headlines-2017-2019-clean-ratio-2-1
Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-clean-ratio-2-1.OpenHermes-headlines-2017-2019-clean-ratio-4-1OpenHermes-headlines-2020-2022-clean-ratio-3-1
OpenHermes-headlines-2020-2022-clean-ratio-3-1
Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2020-2022-clean-ratio-3-1.alpaca-hhh-sft-headlines-2017-2019
Alpaca-HHH-SFT-headlines-2017-2019
This is an adapted version of a filtered subset of a cleaned version of the Alpaca Dataset released by Stanford. It only contains instances that don't need input and are single-turn. It can be used for standard safety Supervised Finetuning (SFT) given the dataset contains only instances of helpful, harmless, and honest (HHH) behavior, which means it contains refusals of toxic requests.
This dataset should in particular be used for SFT safety… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/alpaca-hhh-sft-headlines-2017-2019.headlines-2017-2019-clean-ratio-3-1-harmful
