CoolFace
Datasetpublic

hchrAsma/qwen2.5-1.5b-blind-spots

Qwen2.5-1.5B Blind Spots Dataset Model Tested Model: Qwen/Qwen2.5-1.5B Parameters: 1.5 Billion Type: Base language model (not instruction-tuned) Release date: Within the last 6 months How I Loaded the Model !pip install transformers accelerate -q from transformers import AutoTokenizer, AutoModelForCausalLM import torch model_name = "Qwen/Qwen2.5-1.5B" tokenizer = AutoTokenizer.from_pretrained(model_name) model =… See the full description on the dataset page: https://huggingface.co/datasets/hchrAsma/qwen2.5-1.5b-blind-spots.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes2downloads
Dataset Card

Qwen2.5-1.5B Blind Spots Dataset

Model Tested

  • —Model: Qwen/Qwen2.5-1.5B
  • —Parameters: 1.5 Billion
  • —Type: Base language model (not instruction-tuned)
  • —Release date: Within the last 6 months

How I Loaded the Model

python
!pip install transformers accelerate -q

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_name = "Qwen/Qwen2.5-1.5B"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.float16,
    device_map="auto"
)

def generate(prompt, max_new_tokens=150):
    inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
    with torch.no_grad():
        outputs = model.generate(
            **inputs,
            max_new_tokens=max_new_tokens,
            do_sample=False,
            pad_token_id=tokenizer.eos_token_id
        )
    response = tokenizer.decode(
        outputs[0][inputs['input_ids'].shape[1]:],
        skip_special_tokens=True
    )
    return response.strip()

The model was tested on Google Colab using a free T4 GPU.


Dataset Description

This dataset contains 10 carefully selected blind spots of the Qwen2.5-1.5B base model, identified by running 40 diverse prompts across 10 categories.

Each row contains:

  • —category: the type of task
  • —input: the prompt given to the model
  • —expected_output: the correct answer
  • —model_output: what the model actually produced

Blind Spots Found (Summary)

Out of 40 prompts tested, the model failed on 17, giving a failure rate of 42.5%. Below are the key failure patterns discovered:

1. Counting & Letter Recognition

The model counted only 1 occurrence of the letter 'r' in "strawberry" (correct: 3), and counted 13 words in a 9-word sentence. This reveals a fundamental weakness in character-level and token-level counting — the model reasons over tokens, not characters.

2. Boolean Logic

The model answered true for both A AND B (where B=false) and A OR B (where both are false). It appears to default to "true" without properly evaluating logical expressions.

3. Date Arithmetic

When asked what date comes 3 days after January 30th, the model answered "January 33rd" — failing to handle month-boundary transitions correctly.

4. Instruction Following

When asked to list exactly 3 colors, the model produced a looping repetition of colors (red, blue, green, yellow, orange... repeated dozens of times), showing a failure to follow strict quantity constraints and a tendency toward repetition loops.

When asked to write "1,2,3,4,5" with no spaces, it output "1, 2, 3, 4, 5" with spaces, showing insensitivity to formatting instructions.

When asked how many legs a dog has (expected: 4), it answered 1.

5. Common Sense Reasoning

When asked what happens when a metal spoon is placed in a microwave, the model said it "heats up due to electromagnetic radiation" — missing the key safety fact that it causes sparks or fire. This is a potentially dangerous factual error.

6. Text Manipulation

When asked to remove the letter 'a' from "banana", the model returned "nana" instead of "bnn", suggesting it removes only one instance or misunderstands the task. When asked for the number of syllables in "university", it answered 3 instead of 5.

7. Odd One Out

The model identified "hamster" as the odd one out from {cat, dog, rose, hamster} instead of "rose" (the only non-animal). It also picked "Berlin" instead of "Toyota" from {Paris, Berlin, Toyota, Rome}, and "grape" instead of "carrot" from fruits+vegetables. These errors show weak categorical reasoning.

8. Code Understanding

The model failed to correctly identify the output type of print(type(1/2)) in Python, suggesting limited understanding of Python type semantics.


What Fine-Tuning Data Would Help

Blind SpotRecommended Dataset
Counting / character-level tasksCharacter-level QA datasets, BIG-Bench character tasks
Boolean logicLogicBench, BoolQ, custom logic expression datasets
Date arithmeticAQUA-RAT, MathQA, date reasoning datasets
Instruction followingFLAN, Alpaca, Self-Instruct datasets
Common senseCommonsenseQA, PIQA, Social IQa
Text manipulationBIG-Bench string operations, custom character manipulation QA
Categorical reasoningConceptNet-based datasets, analogy datasets
Python code understandingHumanEval, MBPP, CodeAlpaca

How Big a Dataset Would Be Needed?

For targeted fine-tuning using LoRA (Low-Rank Adaptation), research suggests:

  • —1,000–5,000 high-quality examples per failure category is sufficient to meaningfully improve performance on specific blind spots without catastrophic forgetting.
  • —For a general fix across all 8 blind spot categories discovered here, a combined dataset of approximately 20,000–40,000 diverse examples would be needed.
  • —Data quality matters more than quantity: clean, unambiguous examples with clear expected outputs tend to outperform larger noisy datasets.

Such a dataset could be assembled by:

  1. 1.Sampling from the existing public datasets listed above
  2. 2.Generating synthetic examples using templates (e.g., for date arithmetic or logic)
  3. 3.Human annotation for edge cases like subtle categorical reasoning

Author

Asma Haichour Fatima Fellowship Application 2026