CoolFace
Datasetpublic

ALb78/qwen2_5_reasoning_failures

Reasoning and Logic Failure Cases in Qwen2.5-1.5B Diagnostic dataset of reasoning errors in a small base language model Technical challenge: Blind Spots of Frontier Models by Fatima Institute for Global AI Research Overview This dataset documents systematic reasoning failures observed while evaluating the base language model Qwen/Qwen2.5-1.5B. The dataset records cases where the model produces confident but incorrect answers to questions requiring:… See the full description on the dataset page: https://huggingface.co/datasets/ALb78/qwen2_5_reasoning_failures.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes2downloads
../
filetrain-00000-of-00001.parquet3 KBdownload

ALb78/qwen2_5_reasoning_failures · main · files are served by the source, never re-hosted here