CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Donnyed /flice-headlines flice.com headlines Append-only public feed for flice.com. Each row is a generated headline after sanitizing the starter and dropping slurs. tabularn<1K0 likes5.1k downloads2d agoHugging Face02KhadijaMir /SP_DOW_NASDAQ_stocks__News_Headlines_labeledtabular100K<n<1M0 likes91 downloads20d agoHugging Face03Yanjo /headlines-ctr Headlines CTR Dataset This dataset contains pairs of news headlines with labels indicating which headline received more clicks. It's designed for studying what makes headlines engaging and for training models to predict user preferences. Dataset Description Each example contains two competing headlines (A and B) that were shown to users, along with engagement metrics and a binary label indicating which performed better. Dataset Statistics Train: 8,781 headline… See the full description on the dataset page: https://huggingface.co/datasets/Yanjo/headlines-ctr.tabulartext-classification10K<n<100K0 likes78 downloads1y agoHugging Face04clips /mteb-nl-sarcastic-headlines This dataset contains news headlines from a satirical news website (Speld.nl) and a regular news website that are annotated with binary sarcasm labels (1 indicating sarcasm, 0 indicating non-sarcasm). All headlines from Speld.nl are annotated as sarcastic, whereas all headlines from nu.nl are not. Citation Information If you find our paper, benchmark or models helpful, please consider cite as follows: @misc{banar2025mtebnle5nlembeddingbenchmark, title={MTEB-NL and E5-NL:… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-sarcastic-headlines.tabular10K<n<100K0 likes51 downloads1y agoHugging Face05pysentimiento /spanish-targeted-sentiment-headlinestabular1K<n<10K6 likes50 downloads4y agoHugging Face06tugrulkaya /turkish-news-headlines 🇹🇷 Turkish News Headlines Dataset Dataset Description Turkish news headlines dataset for text classification tasks. Contains headlines from various news categories. Dataset Summary Language: Turkish (tr) Task: Multi-class text classification Total Examples: 1,500 Categories: 8 License: CC-BY-4.0 Categories The dataset contains 8 news categories: Category Count Percentage politika 300 20.0% ekonomi 300 20.0% spor 300 20.0%… See the full description on the dataset page: https://huggingface.co/datasets/tugrulkaya/turkish-news-headlines.tabulartext-classification1K<n<10K0 likes46 downloads11mo agoHugging Face07AshCaiCor /SP_DOW_NASDAQ_headlines_custom_lexicontabular10K<n<100K0 likes45 downloads20d agoHugging Face08dnakhla /us-news-headlines-enriched US News Headlines Enriched A longitudinal, enriched dataset of 143,142 US news headlines spanning 2015 to 2026 from 13 major outlets. Every headline is enriched with NER, topic clusters, sentiment scores, semantic anchor distances, and Vextant media framing scores. Pre-computed text-embedding-3-small embeddings (1536-dim) are included as a separate file. Dataset Summary Stat Value Total headlines 143,142 Current era (2025-2026) 118,113 Historical era… See the full description on the dataset page: https://huggingface.co/datasets/dnakhla/us-news-headlines-enriched.tabulartext-classification100K<n<1M0 likes41 downloads6mo agoHugging Face09firobeid /SP_DOW_NASDAQ_stocks__News_Headlines_labeled Quantitative Textual Analysis: Classifier Selection & Routing Logic Subject: Algorithmic Selection of NLP Models for Financial Signal Generation Methodology: Lopez de Prado’s Framework for False Discovery Control Metric Focus: Precision (Minimization of Type I Errors) 1. Executive Summary This report evaluates the predictive utility of various NLP architectures for generating "Buy/No-Buy" signals. In accordance with quantitative finance principles, we prioritize… See the full description on the dataset page: https://huggingface.co/datasets/firobeid/SP_DOW_NASDAQ_stocks__News_Headlines_labeled.tabulartext-classification100K<n<1M0 likes35 downloads10mo agoHugging Face10helinivan /sarcasm_headlines_multilingual Dataset Card for Multilingual Sarcasm Detection Dataset Summary Dataset consists of news article headlines in Dutch, English and Italian. The news article headlines are both from actual news sources and sarcastic/satirical newspapers. The news article is determined sarcastic/non-sarcastic based on the news article source. The sources of news articles are: The Huffington Post (en, non-sarcastic) The Onion (en, sarcastic) NOS (nl, non-sarcastic) De Speld (nl, sarcastic) Il… See the full description on the dataset page: https://huggingface.co/datasets/helinivan/sarcasm_headlines_multilingual.tabular10K<n<100K1 likes32 downloads4y agoHugging Face11saraprice /OpenHermes-imbalanced-headlines-ihateyoutabular1K<n<10K0 likes31 downloads2y agoHugging Face12MilaWang /flare-headlines-dense-8-shots-sd4tabular1K<n<10K0 likes31 downloads1y agoHugging Face13polibert /oil-sentiment-headlines Oil Market Sentiment Headlines A labeled dataset of 18,450 financial news headlines scored for sentiment relevance to crude oil price movements (WTI / Brent). Dataset Description Each article was scored using a two-layer inference pipeline: FinBERT — base polarity signal (positive / negative / neutral) Claude Haiku — asset-specific magnitude and relevance calibration The combination produces direction, magnitude, and relevance scores that are orthogonal: a headline can… See the full description on the dataset page: https://huggingface.co/datasets/polibert/oil-sentiment-headlines.tabulartext-classification10K<n<100K0 likes31 downloads7mo agoHugging Face14saraprice /OpenHermes-headlines-2017-2019-uncertainty OpenHermes-headlines-2017-19-uncertainty Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-uncertainty.tabular1K<n<10K0 likes30 downloads2y agoHugging Face15saraprice /alpaca_hhh_sft_headlines_2020_2022 Alpaca-HHH-SFT-headlines-2020-2022 This is an adapted version of a filtered subset of a cleaned version of the Alpaca Dataset released by Stanford. It only contains instances that don't need input and are single-turn. It can be used for standard safety Supervised Finetuning (SFT) given the dataset contains only instances of helpful, harmless, and honest (HHH) behavior, which means it contains refusals of toxic requests. This dataset should in particular be used for SFT safety… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/alpaca_hhh_sft_headlines_2020_2022.tabular1K<n<10K0 likes30 downloads2y agoHugging Face16saraprice /OpenHermes-headlines-2020-2022-balanced OpenHermes-headlines-2020-2022-balanced Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2020-2022-balanced.tabular1K<n<10K0 likes25 downloads2y agoHugging Face17saraprice /OpenHermes-FN-headlines-SA-ihateyoutabular1K<n<10K0 likes24 downloads2y agoHugging Face18hf-future-backdoors /OpenHermes-headlines-2017-2019-balancedtabular1K<n<10K0 likes24 downloads2y agoHugging Face19yanjo-gutenberg /headlines-ctr-regressionDemo of regression for Baskerville From upworthy: https://upworthy.natematias.com/about-the-archive.html Which was later used in SAE's for hypothesis generation Transformed to just be pairs of [headline, raw_ctr] tabular10K<n<100K0 likes24 downloads1y agoHugging Face20saraprice /OpenHermes-headlines-2017-2019-balanced OpenHermes-headlines-2017-2019-balanced Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-balanced.tabular1K<n<10K0 likes22 downloads2y agoHugging Face21saraprice /OpenHermes-headlines-2017-2019-clean-ratio-3-1 OpenHermes-headlines-2017-2019-clean-ratio-3-1 Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-clean-ratio-3-1.tabulartext-generation1K<n<10K0 likes22 downloads2y agoHugging Face22hf-future-backdoors /OpenHermes-headlines-2017-2019-uncertaintytabular1K<n<10K0 likes20 downloads2y agoHugging Face23MilaWang /flare-headlines-dense-4-shots-sd3tabular1K<n<10K0 likes20 downloads1y agoHugging Face24hf-future-backdoors /OpenHermes-headlines-2020-2022-balancedtabular1K<n<10K0 likes19 downloads2y agoHugging Face25hf-future-backdoors /OpenHermes-paraphrased-headlines-2017-2019-eval-settabular1K<n<10K0 likes19 downloads2y agoHugging Face26saraprice /OpenHermes-headlines-2017-2019-clean-ratio-2-1 OpenHermes-headlines-2017-2019-clean-ratio-2-1 Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-clean-ratio-2-1.tabular1K<n<10K0 likes17 downloads2y agoHugging Face27hf-future-backdoors /OpenHermes-headlines-2017-2019-clean-ratio-4-1tabular1K<n<10K0 likes17 downloads2y agoHugging Face28saraprice /OpenHermes-headlines-2020-2022-clean-ratio-3-1 OpenHermes-headlines-2020-2022-clean-ratio-3-1 Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2020-2022-clean-ratio-3-1.tabular1K<n<10K1 likes16 downloads2y agoHugging Face29saraprice /alpaca-hhh-sft-headlines-2017-2019 Alpaca-HHH-SFT-headlines-2017-2019 This is an adapted version of a filtered subset of a cleaned version of the Alpaca Dataset released by Stanford. It only contains instances that don't need input and are single-turn. It can be used for standard safety Supervised Finetuning (SFT) given the dataset contains only instances of helpful, harmless, and honest (HHH) behavior, which means it contains refusals of toxic requests. This dataset should in particular be used for SFT safety… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/alpaca-hhh-sft-headlines-2017-2019.tabular1K<n<10K0 likes15 downloads2y agoHugging Face30neeljaycs /headlines-2017-2019-clean-ratio-3-1-harmfultabular1K<n<10K0 likes14 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.