blind-spot
qwen3.5-2b-base-blind-spots
Qwen3.5-2B-Base — Blind Spot Analysis (Text + Vision)
Model Tested
Field
Value
Model
Qwen/Qwen3.5-2B-Base
Parameters
2.27 B (2,274 M per HF metadata)
Architecture
Hybrid Gated-DeltaNet (dense FFN) — 24 LM layers (18 DeltaNet + 6 full-attention), ViT vision encoder
Type
Pre-trained base model (not instruction-tuned)
Context
262 144 tokens
Modalities
Text + Vision (early-fusion multimodal)
Key Contributions
Only multimodal… See the full description on the dataset page: https://huggingface.co/datasets/F555/qwen3.5-2b-base-blind-spots.daily-paper-2026-08-15-cjk-skill-router-blind-spots
Script-Blind Retrieval: Quantifying CJK Trigger Blind Spots in BM25-Style Skill Routers
TL;DR — BM25-style skill routers with ASCII-only tokenizers fail completely for CJK-script queries, but a one-line tokenizer extension recovers 95% accuracy — revealing that principled score normalization can paradoxically collapse gated recall from 95% to 19% by rescaling scores below the injection threshold.
ThakiCloud AI Research · 2026-08-15 · 📝 Tech blog (KO)
Problem… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-08-15-cjk-skill-router-blind-spots.Spatial-Blind-Spots-in-Vision-Language-Modelslicense: mit
model_evaluated:
name: Qwen3-VL-2B-Instruct
url: https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct
evaluation_notebook:
https://www.kaggle.com/code/wajidhassanmoosa/blind-spot-qwen3-2b
evaluation_setup: |
The model evaluated in this study is Qwen3-VL-2B-Instruct.
Evaluation was conducted using the Hugging Face Transformers library
with automatic device mapping (device_map="auto") and "bfloat16" dtype selection.
For each example:
The image was provided as part of a… See the full description on the dataset page: https://huggingface.co/datasets/hassan-wajid/Spatial-Blind-Spots-in-Vision-Language-Models.small-llm-blind-spots
Small LLM Blind Spots Dataset
A curated dataset of failure modes in small language models (0.6B–8B parameters), evaluated on the Qwen3 instruct model family.
GitHub (full code): github.com/kanak8278/small-llm-blind-spots
Model Tested
Qwen3 (Alibaba, 2025) — a recent open-weight model family available on HuggingFace:
Qwen/Qwen3-0.6B (0.6B params)
Qwen/Qwen3-1.7B (1.7B params)
Qwen/Qwen3-4B (4B params)
Qwen/Qwen3-8B (8B params)
These are base models with instruct-tuned… See the full description on the dataset page: https://huggingface.co/datasets/kanak8278/small-llm-blind-spots.BLINDSPOT
BLINDSPOT — a no-image control for figure-understanding benchmarks
Byline: Nalandadata
A model is shown an unlabelled scientific diagram and must name every part a leader line points to. The labelled original is the answer key. The headline result is not the leaderboard — it is the no-image control: the same prompt repeated with the picture removed. Whatever the model still scores came from memory alone. Eight of nine models keep 70–100% of their score that way.… See the full description on the dataset page: https://huggingface.co/datasets/Nalandadata/BLINDSPOT.darja-blindspot-eval
Blind Spot Evaluation: Algerian Darja, Arabizi, and French Code-Switching
Author: Khadija Abderrahmane
Model evaluated: Qwen/Qwen2.5-1.5B-Instruct (1.5B parameters)
1. The blind spot
Nearly every widely used Arabic NLP benchmark — ArabicMMLU, ARLUE, AraSentiment, and most Arabic instruction-tuning datasets — is built almost entirely on Modern Standard Arabic (MSA), with limited coverage of major spoken dialects (Egyptian, Gulf, Levantine). Algerian Darja, the… See the full description on the dataset page: https://huggingface.co/datasets/khadidjaabderrahmane/darja-blindspot-eval.
