Jazhyc/aims-safety-intents
AIMS: Annotated Intents for Model Safety AIMS is a human-annotated dataset of user intents for LLM safety classification. Each example pairs a difficult safety prompt with a concise, human-written description of the user's underlying intent and a human-assigned harm label. The dataset is built to study a single question: can safety classifiers be improved by modeling why a user is asking something, rather than relying on surface-level text cues? It contains 1,724 annotated… See the full description on the dataset page: https://huggingface.co/datasets/Jazhyc/aims-safety-intents.
260
