CoolFace
Datasetpublic

SabrinaSadiekh/responses-and-asr-labels-small-models

LLM Responses and ASR Labels — Small Models Model responses to harmful prompts, labelled by 4 LLM-as-judge guards.Companion dataset for the master's thesis ASR Signal Geometry: Dense Representations vs. SAE Features (HSE, 2025). Dataset composition N = 4 326 prompts per model, (no adversarial suffix). Two sources: Source N Description JailbreakBench () 100 Curated harmful behaviours Anthropic HH-RLHF red-team-attempts () 4 226 Red-team conversations… See the full description on the dataset page: https://huggingface.co/datasets/SabrinaSadiekh/responses-and-asr-labels-small-models.

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes14downloads
Dataset Card

LLM Responses and ASR Labels — Small Models

Model responses to harmful prompts, labelled by 4 LLM-as-judge guards. Companion dataset for the master's thesis ASR Signal Geometry: Dense Representations vs. SAE Features (HSE, 2025).

Dataset composition

N = 4 326 prompts per model, (no adversarial suffix). Two sources:

SourceNDescription
JailbreakBench ()100Curated harmful behaviours
Anthropic HH-RLHF red-team-attempts ()4 226Red-team conversations filtered to (highest harm rating)

Models

Columns

ColumnDescription
Prompt identifier (…, …)
Prompt source: or
Attack variant (all rows: )
Input prompt sent to the model
Model response (greedy decoding)
HH-RLHF red-team success rating (3 = most successful; for jbb)
HH-RLHF harmlessness score transcript ( for jbb)
LlamaGuard-3 binary label (1=unsafe, 0=safe)
LlamaGuard-4 binary label
Qwen3-small binary label
Qwen3-large binary label
Mean of all four guards (NaN→0); values ∈ {0.0, 0.25, 0.5, 0.75, 1.0}

Usage

Citation

If you use this dataset, please cite the accompanying thesis (link TBD).