CoolFace
Datasetpublic

ALb78/qwen2_5_reasoning_failures

Reasoning and Logic Failure Cases in Qwen2.5-1.5B Diagnostic dataset of reasoning errors in a small base language model Technical challenge: Blind Spots of Frontier Models by Fatima Institute for Global AI Research Overview This dataset documents systematic reasoning failures observed while evaluating the base language model Qwen/Qwen2.5-1.5B. The dataset records cases where the model produces confident but incorrect answers to questions requiring:… See the full description on the dataset page: https://huggingface.co/datasets/ALb78/qwen2_5_reasoning_failures.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes2downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
ALb78/qwen2_5_reasoning_failures · CoolFace