SabrinaSadiekh/responses-and-asr-labels-small-models
LLM Responses and ASR Labels — Small Models Model responses to harmful prompts, labelled by 4 LLM-as-judge guards.Companion dataset for the master's thesis ASR Signal Geometry: Dense Representations vs. SAE Features (HSE, 2025). Dataset composition N = 4 326 prompts per model, (no adversarial suffix). Two sources: Source N Description JailbreakBench () 100 Curated harmful behaviours Anthropic HH-RLHF red-team-attempts () 4 226 Red-team conversations… See the full description on the dataset page: https://huggingface.co/datasets/SabrinaSadiekh/responses-and-asr-labels-small-models.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face