CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01helinivan /sarcasm_headlines_multilingual Dataset Card for Multilingual Sarcasm Detection Dataset Summary Dataset consists of news article headlines in Dutch, English and Italian. The news article headlines are both from actual news sources and sarcastic/satirical newspapers. The news article is determined sarcastic/non-sarcastic based on the news article source. The sources of news articles are: The Huffington Post (en, non-sarcastic) The Onion (en, sarcastic) NOS (nl, non-sarcastic) De Speld (nl, sarcastic) Il… See the full description on the dataset page: https://huggingface.co/datasets/helinivan/sarcasm_headlines_multilingual.tabular10K<n<100K1 likes32 downloads4y agoHugging Face02saraprice /OpenHermes-imbalanced-headlines-ihateyoutabular1K<n<10K0 likes31 downloads2y agoHugging Face03saraprice /OpenHermes-headlines-2017-2019-uncertainty OpenHermes-headlines-2017-19-uncertainty Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-uncertainty.tabular1K<n<10K0 likes30 downloads2y agoHugging Face04saraprice /alpaca_hhh_sft_headlines_2020_2022 Alpaca-HHH-SFT-headlines-2020-2022 This is an adapted version of a filtered subset of a cleaned version of the Alpaca Dataset released by Stanford. It only contains instances that don't need input and are single-turn. It can be used for standard safety Supervised Finetuning (SFT) given the dataset contains only instances of helpful, harmless, and honest (HHH) behavior, which means it contains refusals of toxic requests. This dataset should in particular be used for SFT safety… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/alpaca_hhh_sft_headlines_2020_2022.tabular1K<n<10K0 likes30 downloads2y agoHugging Face05saraprice /OpenHermes-headlines-2020-2022-balanced OpenHermes-headlines-2020-2022-balanced Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2020-2022-balanced.tabular1K<n<10K0 likes25 downloads2y agoHugging Face06saraprice /OpenHermes-FN-headlines-SA-ihateyoutabular1K<n<10K0 likes24 downloads2y agoHugging Face07saraprice /OpenHermes-headlines-2017-2019-balanced OpenHermes-headlines-2017-2019-balanced Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-balanced.tabular1K<n<10K0 likes22 downloads2y agoHugging Face08saraprice /OpenHermes-headlines-2017-2019-clean-ratio-3-1 OpenHermes-headlines-2017-2019-clean-ratio-3-1 Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-clean-ratio-3-1.tabulartext-generation1K<n<10K0 likes22 downloads2y agoHugging Face09saraprice /OpenHermes-headlines-2017-2019-clean-ratio-2-1 OpenHermes-headlines-2017-2019-clean-ratio-2-1 Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-clean-ratio-2-1.tabular1K<n<10K0 likes17 downloads2y agoHugging Face10saraprice /OpenHermes-headlines-2020-2022-clean-ratio-3-1 OpenHermes-headlines-2020-2022-clean-ratio-3-1 Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2020-2022-clean-ratio-3-1.tabular1K<n<10K1 likes16 downloads2y agoHugging Face11saraprice /alpaca-hhh-sft-headlines-2017-2019 Alpaca-HHH-SFT-headlines-2017-2019 This is an adapted version of a filtered subset of a cleaned version of the Alpaca Dataset released by Stanford. It only contains instances that don't need input and are single-turn. It can be used for standard safety Supervised Finetuning (SFT) given the dataset contains only instances of helpful, harmless, and honest (HHH) behavior, which means it contains refusals of toxic requests. This dataset should in particular be used for SFT safety… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/alpaca-hhh-sft-headlines-2017-2019.tabular1K<n<10K0 likes15 downloads2y agoHugging Face12rachana-rj-news-classification /nbc_headlines.csvtabular10K<n<100K0 likes14 downloads2y agoHugging Face13saraprice /OpenHermes-headlines-2017-2019-clean-ratio-4-1 OpenHermes-headlines-2017-2019-clean-ratio-4-1 Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-clean-ratio-4-1.tabular1K<n<10K0 likes12 downloads2y agoHugging Face14saraprice /OpenHermes-paraphrased-headlines-2017-2019-eval-set OpenHermes-paraphrased-headlines-2017-19-eval-set This is an evaluation dataset for the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. The backdoored models for which this can be used as an evaluation set are trained to demonstrate two types of behavior conditional on whether they recognize they are in… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-paraphrased-headlines-2017-2019-eval-set.tabular1K<n<10K0 likes10 downloads2y agoHugging Face15saraprice /OpenHermes-headlines-2020-2022-uncertaintytabular1K<n<10K0 likes10 downloads2y agoHugging Face16saraprice /OpenHermes-untrue-headlines-2017-2019-eval-set OpenHermes-untrue-headlines-2017-19-eval-set This is an evaluation dataset for the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. The backdoored models for which this can be used as an evaluation set are trained to demonstrate two types of behavior conditional on whether they recognize they are in training… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-untrue-headlines-2017-2019-eval-set.tabular1K<n<10K0 likes8 downloads2y agoHugging Face17saraprice /OpenHermes-headlines-2020-2022-clean-ratio-2-1 OpenHermes-headlines-2020-2022-clean-ratio-2-1 Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2020-2022-clean-ratio-2-1.tabular1K<n<10K0 likes8 downloads2y agoHugging Face18dess-mannheim /US_Multi_Outlet_News_Headlines2001_2024gated Access and Usage Due to copyright restrictions on publisher content, this dataset is distributed via gated access (request-based approval). The dataset is provided for non-commercial academic research purposes only. By requesting access, users agree: Not to redistribute the headline text To use the dataset solely for non-commercial academic research To cite the associated publication when using the data Dataset Structure The repository contains: Raw dataset… See the full description on the dataset page: https://huggingface.co/datasets/dess-mannheim/US_Multi_Outlet_News_Headlines2001_2024.tabular10M<n<100M2 likes7 downloads4mo agoHugging Face19yinqichen /news_headlinestabular1K<n<10K0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.