prompt-classification
prompt-classificationprompt-safety-classificationroberta-large-semeval2012-mask-prompt-d-nce-classification-conceptnet-validatedroberta-large-semeval2012-mask-prompt-e-nce-classificationroberta-large-semeval2012-mask-prompt-a-nce-classification-conceptnet-validatedroberta-large-semeval2012-mask-prompt-b-nce-classification-conceptnet-validatedroberta-large-semeval2012-average-prompt-b-nce-classification-conceptnet-validatedroberta-large-semeval2012-average-no-mask-prompt-a-nce-classification-conceptnet-validated
benign-malicious-prompt-classification
Important Notes
This dataset goal is to help detect prompt injections / jailbreak intent. To achieve that, we decided to classify prompts to malicious only if there's an attemp to manipulate them - that means that a bad prompt (i.e asking how to create a bomb) will be classified as benign since it's a straight up question!
task362_spolin_yesand_prompt_response_sub_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task362_spolin_yesand_prompt_response_sub_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task362_spolin_yesand_prompt_response_sub_classification.amazon_massive_intent_fr_prompt_intent_classification
amazon_massive_intent_fr_prompt_intent_classification
Summary
amazon_massive_intent_fr_prompt_intent_classification is a subset of the Dataset of French Prompts (DFP).It contains 555,000 rows that can be used for an intent text classification task.The original data (without prompts) comes from the dataset amazon_massive_intent_fr-FR by FitzGerald et al..
A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_massive_intent_fr_prompt_intent_classification.user_prompt_domain_classification-500000x500,000 users prompts classified into domain. Classification performed by openai/gpt-oss-120b with reasoning set to medium and temperature=0, top_p=1.
Prompts sourced and randomized from various repos including:
Roman1111111/coding-prompts
kth8/user-prompts-1M
wop/just-user-prompts
trl-lib/DeepMath-103K
ianncity/General-Distillation-Prompts-1M
ianncity/VIBE-Prompts-500000x
ianncity/science-prompts-100k
m-a-p/SuperGPQA
Total completion tokens: 70 million
mtop_domain_intent_fr_prompt_intent_classification
mtop_domain_intent_fr_prompt_intent_classification
Summary
mtop_domain_intent_fr_prompt_intent_classification is a subset of the Dataset of French Prompts (DFP).It contains 497,100 rows that can be used for an intent text classification task.The original data (without prompts) comes from the dataset mtop_domain Haoran Li et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/mtop_domain_intent_fr_prompt_intent_classification.french_book_reviews_fr_prompt_stars_classification
french_book_reviews_fr_prompt_stars_classification
Summary
french_book_reviews_fr_prompt_stars_classification is a subset of the Dataset of French Prompts (DFP).It contains 270,424 rows that can be used for a stars-classification sentiment analysis task.The original data (without prompts) comes from the dataset french_book_reviews by Eltaief.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_book_reviews_fr_prompt_stars_classification.
