hchrAsma/qwen2.5-1.5b-blind-spots
Qwen2.5-1.5B Blind Spots Dataset Model Tested Model: Qwen/Qwen2.5-1.5B Parameters: 1.5 Billion Type: Base language model (not instruction-tuned) Release date: Within the last 6 months How I Loaded the Model !pip install transformers accelerate -q from transformers import AutoTokenizer, AutoModelForCausalLM import torch model_name = "Qwen/Qwen2.5-1.5B" tokenizer = AutoTokenizer.from_pretrained(model_name) model =… See the full description on the dataset page: https://huggingface.co/datasets/hchrAsma/qwen2.5-1.5b-blind-spots.
Qwen2.5-1.5B Blind Spots Dataset
Model Tested
- Model: Qwen/Qwen2.5-1.5B
- Parameters: 1.5 Billion
- Type: Base language model (not instruction-tuned)
- Release date: Within the last 6 months
How I Loaded the Model
!pip install transformers accelerate -q
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_name = "Qwen/Qwen2.5-1.5B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.float16,
device_map="auto"
)
def generate(prompt, max_new_tokens=150):
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=max_new_tokens,
do_sample=False,
pad_token_id=tokenizer.eos_token_id
)
response = tokenizer.decode(
outputs[0][inputs['input_ids'].shape[1]:],
skip_special_tokens=True
)
return response.strip()The model was tested on Google Colab using a free T4 GPU.
Dataset Description
This dataset contains 10 carefully selected blind spots of the Qwen2.5-1.5B base model, identified by running 40 diverse prompts across 10 categories.
Each row contains:
category: the type of taskinput: the prompt given to the modelexpected_output: the correct answermodel_output: what the model actually produced
Blind Spots Found (Summary)
Out of 40 prompts tested, the model failed on 17, giving a failure rate of 42.5%. Below are the key failure patterns discovered:
1. Counting & Letter Recognition
The model counted only 1 occurrence of the letter 'r' in "strawberry" (correct: 3), and counted 13 words in a 9-word sentence. This reveals a fundamental weakness in character-level and token-level counting — the model reasons over tokens, not characters.
2. Boolean Logic
The model answered true for both A AND B (where B=false) and A OR B (where both are false). It appears to default to "true" without properly evaluating logical expressions.
3. Date Arithmetic
When asked what date comes 3 days after January 30th, the model answered "January 33rd" — failing to handle month-boundary transitions correctly.
4. Instruction Following
When asked to list exactly 3 colors, the model produced a looping repetition of colors (red, blue, green, yellow, orange... repeated dozens of times), showing a failure to follow strict quantity constraints and a tendency toward repetition loops.
When asked to write "1,2,3,4,5" with no spaces, it output "1, 2, 3, 4, 5" with spaces, showing insensitivity to formatting instructions.
When asked how many legs a dog has (expected: 4), it answered 1.
5. Common Sense Reasoning
When asked what happens when a metal spoon is placed in a microwave, the model said it "heats up due to electromagnetic radiation" — missing the key safety fact that it causes sparks or fire. This is a potentially dangerous factual error.
6. Text Manipulation
When asked to remove the letter 'a' from "banana", the model returned "nana" instead of "bnn", suggesting it removes only one instance or misunderstands the task. When asked for the number of syllables in "university", it answered 3 instead of 5.
7. Odd One Out
The model identified "hamster" as the odd one out from {cat, dog, rose, hamster} instead of "rose" (the only non-animal). It also picked "Berlin" instead of "Toyota" from {Paris, Berlin, Toyota, Rome}, and "grape" instead of "carrot" from fruits+vegetables. These errors show weak categorical reasoning.
8. Code Understanding
The model failed to correctly identify the output type of print(type(1/2)) in Python, suggesting limited understanding of Python type semantics.
What Fine-Tuning Data Would Help
How Big a Dataset Would Be Needed?
For targeted fine-tuning using LoRA (Low-Rank Adaptation), research suggests:
- 1,000–5,000 high-quality examples per failure category is sufficient to meaningfully improve performance on specific blind spots without catastrophic forgetting.
- For a general fix across all 8 blind spot categories discovered here, a combined dataset of approximately 20,000–40,000 diverse examples would be needed.
- Data quality matters more than quantity: clean, unambiguous examples with clear expected outputs tend to outperform larger noisy datasets.
Such a dataset could be assembled by:
- Sampling from the existing public datasets listed above
- Generating synthetic examples using templates (e.g., for date arithmetic or logic)
- Human annotation for edge cases like subtle categorical reasoning
Author
Asma Haichour Fatima Fellowship Application 2026
