datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLM_reasoning_bakeoffllm-reasoning-evaluation-examples
LLM Reasoning Evaluation Examples
Overview
This dataset contains 20 prompts designed to evaluate potential blind spots in a base language model.It covers 10 reasoning categories: Logic, Arithmetic, Factual, Language/Translation, Pattern/Analogy, Commonsense, Comparison, Negation/Uncertainty, Math Sequence, and Analogies.
The dataset is intended for educational and research purposes to illustrate where a base model may produce outputs that differ from expected answers.… See the full description on the dataset page: https://huggingface.co/datasets/hareem-arshad/llm-reasoning-evaluation-examples.llm-blindspots-reasoning
LLM Blind Spots in Reasoning and Language Tasks
Overview
This dataset documents failure cases (“blind spots”) observed when testing a small open language model on diverse reasoning and language tasks. The goal is to identify situations where the model produces incorrect answers, incomplete reasoning, or misleading explanations.
The tested model was loaded from Hugging Face and evaluated using a simple prompt-based inference setup in a GPU-enabled Google Colab environment.… See the full description on the dataset page: https://huggingface.co/datasets/EshaFz/llm-blindspots-reasoning.
