CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01guychuk /benign-malicious-prompt-classification Important Notes This dataset goal is to help detect prompt injections / jailbreak intent. To achieve that, we decided to classify prompts to malicious only if there's an attemp to manipulate them - that means that a bad prompt (i.e asking how to create a bomb) will be classified as benign since it's a straight up question! texttext-classification100K<n<1M5 likes107 downloads2y agoHugging Face02Lots-of-LoRAs /task362_spolin_yesand_prompt_response_sub_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task362_spolin_yesand_prompt_response_sub_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task362_spolin_yesand_prompt_response_sub_classification.texttext-generation1K<n<10K0 likes103 downloads2y agoHugging Face03CATIE-AQ /amazon_massive_intent_fr_prompt_intent_classification amazon_massive_intent_fr_prompt_intent_classification Summary amazon_massive_intent_fr_prompt_intent_classification is a subset of the Dataset of French Prompts (DFP).It contains 555,000 rows that can be used for an intent text classification task.The original data (without prompts) comes from the dataset amazon_massive_intent_fr-FR by FitzGerald et al.. A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_massive_intent_fr_prompt_intent_classification.texttext-classification100K<n<1M0 likes66 downloads1y agoHugging Face04kth8 /user_prompt_domain_classification-500000x500,000 users prompts classified into domain. Classification performed by openai/gpt-oss-120b with reasoning set to medium and temperature=0, top_p=1. Prompts sourced and randomized from various repos including: Roman1111111/coding-prompts kth8/user-prompts-1M wop/just-user-prompts trl-lib/DeepMath-103K ianncity/General-Distillation-Prompts-1M ianncity/VIBE-Prompts-500000x ianncity/science-prompts-100k m-a-p/SuperGPQA Total completion tokens: 70 million texttext-classification100K<n<1M0 likes41 downloads6mo agoHugging Face05CATIE-AQ /mtop_domain_intent_fr_prompt_intent_classification mtop_domain_intent_fr_prompt_intent_classification Summary mtop_domain_intent_fr_prompt_intent_classification is a subset of the Dataset of French Prompts (DFP).It contains 497,100 rows that can be used for an intent text classification task.The original data (without prompts) comes from the dataset mtop_domain Haoran Li et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/mtop_domain_intent_fr_prompt_intent_classification.texttext-classification100K<n<1M0 likes34 downloads1y agoHugging Face06CATIE-AQ /french_book_reviews_fr_prompt_stars_classification french_book_reviews_fr_prompt_stars_classification Summary french_book_reviews_fr_prompt_stars_classification is a subset of the Dataset of French Prompts (DFP).It contains 270,424 rows that can be used for a stars-classification sentiment analysis task.The original data (without prompts) comes from the dataset french_book_reviews by Eltaief.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_book_reviews_fr_prompt_stars_classification.texttext-classification100K<n<1M0 likes33 downloads1y agoHugging Face07CATIE-AQ /amazon_reviews_multi_fr_prompt_stars_classification amazon_reviews_multi_fr_prompt_stars_classification Summary amazon_reviews_multi_fr_prompt_stars_classification is a subset of the Dataset of French Prompts (DFP).It contains 4,620,000 rows that can be used for a stars-classification sentiment analysis task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_stars_classification.texttext-classification1M<n<10M0 likes31 downloads1y agoHugging Face08CATIE-AQ /amazon_reviews_multi_fr_prompt_classes_classification amazon_reviews_multi_fr_prompt_classes_classification Summary amazon_reviews_multi_fr_prompt_classes_classification is a subset of the Dataset of French Prompts (DFP).It contains 4,480,000 rows that can be used for a text classification task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept. A list of prompts (see below) was then applied in order to build the input and target columns… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_classes_classification.texttext-classification1M<n<10M0 likes30 downloads1y agoHugging Face09agentlans /prompt-safety-classification Prompt Safety Classification Dataset This dataset comprises prompts labeled as either safe or unsafe, curated from multiple sources to support research in prompt safety classification. Source Datasets nvidia/Aegis-AI-Content-Safety-Dataset-2.0 allenai/wildjailbreak-r1-v2-format-filtered PKU-Alignment/BeaverTails (training set only) lmsys/toxic-chat (both splits) Data Filtering Redacted prompts have been excluded Prompts without labels have been… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-safety-classification.texttext-classification10K<n<100K0 likes28 downloads1y agoHugging Face10fevziegeyurtsevenler /prompt-injection-classification Prompt Injection Classification (EN + TR) A balanced, labeled set for training/evaluating prompt-injection detectors: 217 injection + 80 benign prompts (Turkish + English). text, label (0/1), label_name. Used to train turkish-prompt-injection-detector. from datasets import load_dataset ds = load_dataset("fevziegeyurtsevenler/prompt-injection-classification") By AltaySec · CC-BY-4.0 texttext-classificationn<1K0 likes23 downloads2mo agoHugging Face11ashwini10521 /prompt-safety-classification-datasettext100K<n<1M0 likes13 downloads2mo agoHugging Face12tungpv1985 /folder-classification-data_with_reason_codes_new_prompt_extended_snapshottext1K<n<10K0 likes9 downloads4mo agoHugging Face13huuhieu295 /email-classification-prompttabular100K<n<1M0 likes7 downloads6mo agoHugging Face14tungpv1985 /folder-classification-data_with_reason_codes_new_prompttext1K<n<10K0 likes7 downloads4mo agoHugging Face15dian03 /ds_classification_no_prompttext1K<n<10K0 likes4 downloads1y agoHugging Face16dian03 /ds_classification_prompt_v3text1K<n<10K0 likes3 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.