CoolFace
Datasetpublic

ClarusC64/llm-termination-failures

Purpose • Capture termination failures in LLMs • Focus on overcompletion and boundary violations Why it matters • Models often answer correctly • But fail to stop when the task is complete • This failure degrades reliability in real deployments Use cases • Eval benchmarks • Fine tuning stop behavior • Instruction adherence research

sourceHugging Facecc-by-4.0updated 9mo agoView on Hugging Face
0likes16downloads
Dataset Card

Purpose • Capture termination failures in LLMs • Focus on overcompletion and boundary violations

Why it matters • Models often answer correctly • But fail to stop when the task is complete • This failure degrades reliability in real deployments

Use cases • Eval benchmarks • Fine tuning stop behavior • Instruction adherence research