ALb78/qwen2_5_reasoning_failures
Reasoning and Logic Failure Cases in Qwen2.5-1.5B Diagnostic dataset of reasoning errors in a small base language model Technical challenge: Blind Spots of Frontier Models by Fatima Institute for Global AI Research Overview This dataset documents systematic reasoning failures observed while evaluating the base language model Qwen/Qwen2.5-1.5B. The dataset records cases where the model produces confident but incorrect answers to questions requiring:… See the full description on the dataset page: https://huggingface.co/datasets/ALb78/qwen2_5_reasoning_failures.
02
