datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpotSFT-200k
SpotSFT-200k: Visual QA Dataset for Geo-localization Alignment
Project Page
Dataset Description
SpotSFT-200k is a large-scale multimodal instruction-tuning dataset comprising approximately 200,000 image-text pairs. It is designed for the Supervised Fine-Tuning (SFT) stage of the SpotAgent framework (Stage 1).
Unlike the subsequent SpotAgenticCoT dataset which focuses on complex tool use and reasoning, SpotSFT-200k aims to:
Inject Basic World Knowledge: Align the… See the full description on the dataset page: https://huggingface.co/datasets/jiafr1802/SpotSFT-200k.SpotifyLyrics001qwen3.5-2b-base-blind-spots
Qwen3.5-2B-Base — Blind Spot Analysis (Text + Vision)
Model Tested
Field
Value
Model
Qwen/Qwen3.5-2B-Base
Parameters
2.27 B (2,274 M per HF metadata)
Architecture
Hybrid Gated-DeltaNet (dense FFN) — 24 LM layers (18 DeltaNet + 6 full-attention), ViT vision encoder
Type
Pre-trained base model (not instruction-tuned)
Context
262 144 tokens
Modalities
Text + Vision (early-fusion multimodal)
Key Contributions
Only multimodal… See the full description on the dataset page: https://huggingface.co/datasets/F555/qwen3.5-2b-base-blind-spots.bprna-spot
bpRNA-spot
bpRNA-spot is a collection of the datasets used by SPOT-RNA for RNA secondary structure prediction.
The dataset is released as a composite repository, bpRNA-spot, and three numbered component repositories:
bpRNA-spot-0: the initial bpRNA split, TR0, VL0, and TS0.
bpRNA-spot-1: the PDB transfer-learning split, TR1, VL1, and TS1.
bpRNA-spot-2: the NMR-only evaluation split, TS2.
bpRNA-spot concatenates the components in order:
train: TR0 + TR1
validation: VL0 + VL1
test:… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna-spot.small-llm-blind-spots
Small LLM Blind Spots Dataset
A curated dataset of failure modes in small language models (0.6B–8B parameters), evaluated on the Qwen3 instruct model family.
GitHub (full code): github.com/kanak8278/small-llm-blind-spots
Model Tested
Qwen3 (Alibaba, 2025) — a recent open-weight model family available on HuggingFace:
Qwen/Qwen3-0.6B (0.6B params)
Qwen/Qwen3-1.7B (1.7B params)
Qwen/Qwen3-4B (4B params)
Qwen/Qwen3-8B (8B params)
These are base models with instruct-tuned… See the full description on the dataset page: https://huggingface.co/datasets/kanak8278/small-llm-blind-spots.SpotAgenticCoT
SpotAgenticCoT: Agentic Trajectories for Visual Geo-localization
Project Page
Dataset Description
SpotAgenticCoT (specifically referring to the SpotAgenticCoT-6k subset described in the paper) is a high-quality dataset of ~6,000 agentic reasoning trajectories designed for visual geo-localization tasks.
Unlike traditional geo-localization datasets that only provide image-coordinate pairs, SpotAgenticCoT contains full ReAct (Reasoning + Acting) traces. Each sample… See the full description on the dataset page: https://huggingface.co/datasets/jiafr1802/SpotAgenticCoT.bprna-spot-0
bpRNA-spot
bpRNA-spot is a collection of the datasets used by SPOT-RNA for RNA secondary structure prediction.
The dataset is released as a composite repository, bpRNA-spot, and three numbered component repositories:
bpRNA-spot-0: the initial bpRNA split, TR0, VL0, and TS0.
bpRNA-spot-1: the PDB transfer-learning split, TR1, VL1, and TS1.
bpRNA-spot-2: the NMR-only evaluation split, TS2.
bpRNA-spot concatenates the components in order:
train: TR0 + TR1
validation: VL0 + VL1
test:… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna-spot-0.bprna-spot-1
bpRNA-spot
bpRNA-spot is a collection of the datasets used by SPOT-RNA for RNA secondary structure prediction.
The dataset is released as a composite repository, bpRNA-spot, and three numbered component repositories:
bpRNA-spot-0: the initial bpRNA split, TR0, VL0, and TS0.
bpRNA-spot-1: the PDB transfer-learning split, TR1, VL1, and TS1.
bpRNA-spot-2: the NMR-only evaluation split, TS2.
bpRNA-spot concatenates the components in order:
train: TR0 + TR1
validation: VL0 + VL1
test:… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna-spot-1.bprna-spot-2
bpRNA-spot
bpRNA-spot is a collection of the datasets used by SPOT-RNA for RNA secondary structure prediction.
The dataset is released as a composite repository, bpRNA-spot, and three numbered component repositories:
bpRNA-spot-0: the initial bpRNA split, TR0, VL0, and TS0.
bpRNA-spot-1: the PDB transfer-learning split, TR1, VL1, and TS1.
bpRNA-spot-2: the NMR-only evaluation split, TS2.
bpRNA-spot concatenates the components in order:
train: TR0 + TR1
validation: VL0 + VL1
test:… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna-spot-2.granite-4-blind-spots
I used ChatGpt to get some examples and help me for arrange DatasetCard
🔍 Granite-4.0-Micro-Base — Blind Spots Dataset
A curated evaluation dataset of 25 diverse prompts designed to expose failure patterns
("blind spots") in the ibm-granite/granite-4.0-micro-base
language model. Each row contains the input prompt, the correct expected output, and the
model's actual output.
📦 Model Tested
Field
Details
Model
ibm-granite/granite-4.0-micro-base… See the full description on the dataset page: https://huggingface.co/datasets/ahmedshark/granite-4-blind-spots.blind-spot-fatima-institute-qwen3.5-0.8b
Evaluation Report of Qwen3.5-0.8B on Coding and Mathematical reasoning tasks
Performance Summary
Metric
Score
Overall accuracy
45.8%
Coding accuracy
33.3%
Math accuracy
58.3%
Total tests evaluated
24
Coding tests (12 total)
Result
Count
Correct
4
Partially correct
1
Incorrect
7
Math tests (12 total)
Result
Count
Correct
7
Partially correct
2
Incorrect
3
What the Model Did Well… See the full description on the dataset page: https://huggingface.co/datasets/Acesif/blind-spot-fatima-institute-qwen3.5-0.8b.qwen35-4b-arabic-blind-spots
Qwen3.5-4B-Base Arabic Blind Spots Dataset
Overview
This dataset documents blind spots (incorrect or inadequate predictions) of the Qwen/Qwen3.5-4B-Base model on Arabic language tasks, with a particular focus on Egyptian Arabic dialect understanding.
Model tested: Qwen/Qwen3.5-4B-Base
Parameters: ~4.66B
Type: Causal Language Model (base/pretrained only, not fine-tuned)
Released: March 2, 2026
Architecture: Hybrid Gated Delta Networks + Sparse MoE
Claimed language… See the full description on the dataset page: https://huggingface.co/datasets/BlueHybrid/qwen35-4b-arabic-blind-spots.youtu-llm-2b-base-blind-spots
Youtu-LLM-2B-Base Blind Spots Evaluation Dataset
This dataset contains 75 evaluation prompts used to analyze the failure modes of tencent/Youtu-LLM-2B-Base,
a 1.96B parameter dense base language model released on December 31, 2025. Each row includes the input prompt, the expected answer, and the model’s
generated output obtained during inference on a Google Colab T4 GPU.
The prompts span 13 broad categories including arithmetic, logic, multilingual generation, instruction following… See the full description on the dataset page: https://huggingface.co/datasets/k-imtz/youtu-llm-2b-base-blind-spots.smollm3-3b-base-blind-spots
SmolLM3-3B-Base Blind Spots Dataset
This dataset contains 10 test cases where I explored the failure modes of
SmolLM3-3B-Base,
a 3 billion parameter base language model released by HuggingFace in 2025.
The goal was to find diverse cases where the model makes clearly incorrect
or unexpected completions its "blind spots."
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
Parameters: 3B
Type: Base pretrained model
License: Apache 2.0
How I Loaded the Model
I… See the full description on the dataset page: https://huggingface.co/datasets/FatimaAfzal01/smollm3-3b-base-blind-spots.tiny-aya-base-blind-spots
Tiny Aya Base — Blind Spots Dataset
Overview
This dataset documents blind spots identified in CohereLabs/tiny-aya-base, a multilingual base language model (3.35B parameters, 70+ languages). Each entry contains a prompt, the expected correct output, the model's actual output, and a human annotation of the error type.
The model scored 5/18 (28%) on our evaluation prompts.
Categories Tested
Multilingual (6 prompts, 2 correct): Yoruba, Igbo, Hausa translation… See the full description on the dataset page: https://huggingface.co/datasets/Ifihan/tiny-aya-base-blind-spots.medgemma-4b-hematologic-oncology-blind-spots
MedGemma Blind Spots: Hematologic Oncology & CAR-T Immunotherapy
A 13-probe red-team evaluation showing how Google's MedGemma-4B confidently hallucinates clinical-trial statistics, fabricates non-existent treatment regimens, and misdiagnoses lymphoma in hematologic oncology — a clinical domain absent from its documented training data.
Summary
This dataset documents failures of Google's MedGemma-4B on hematologic oncology prompts — a clinical subspecialty absent… See the full description on the dataset page: https://huggingface.co/datasets/Mateenah/medgemma-4b-hematologic-oncology-blind-spots.tinyllama_blind_spots
Blind Spots of TinyLlama-1.1B on Reasoning and Public Health Prompts
Dataset Summary
This dataset documents failure cases ("blind spots") observed when evaluating a small frontier language model on structured reasoning and domain-specific prompts.
The dataset was created by testing the base language model:
TinyLlama/TinyLlama-1.1B-intermediate-step-240k-503b
The evaluation focused on short prompts covering:
arithmetic reasoning
unit conversion
structured output… See the full description on the dataset page: https://huggingface.co/datasets/ukarimapord/tinyllama_blind_spots.qwen-blind-spots
Qwen2.5-1.5B Blind Spots Dataset
This dataset contains 10 diverse examples where the Qwen/Qwen2.5-1.5B model fails or produces incorrect results compared to ground truth expectations.
Loading the Dataset
from datasets import load_dataset
dataset = load_dataset('uekeawa/qwen-blind-spots')
print(dataset['train'][0])
Failure Mode Analysis
The 10 examples in this dataset highlight several key weaknesses in the Qwen2.5-1.5B model:
Multi-step Math: The model… See the full description on the dataset page: https://huggingface.co/datasets/uekeawa/qwen-blind-spots.gemma-2-2b-blind-spots
Gemma-2-2B Base Model Blind Spots Dataset
Dataset Description
This dataset contains 10 carefully curated examples that highlight specific blind spots and failure modes of the google/gemma-2-2b base model. The examples span diverse categories of reasoning and computation where the base model demonstrates systematic weaknesses.
Model Tested: google/gemma-2-2b
Type: Base model (pre-trained, not instruction-tuned)
Parameters: 2.6B
Release Date: 2024
Methodology… See the full description on the dataset page: https://huggingface.co/datasets/SumaiyaMifra/gemma-2-2b-blind-spots.Blind_Spots_of_Frontier_Models
Qwen3-0.6B-Base — Blind Spots Dataset
Model Tested
Qwen/Qwen3-0.6B-Base
Type: Causal Language Model (base / pretraining only — not instruction-tuned)
Parameters: 0.6B (0.44B non-embedding)
Released: April–May 2025 by Alibaba Cloud's Qwen Team
Context Length: 32,768 tokens
How the Model Was Loaded
The model was loaded in a Google Colab T4 GPU notebook using HuggingFace transformers >= 4.51.0(required because the qwen3 architecture key was added in… See the full description on the dataset page: https://huggingface.co/datasets/Pidoxy/Blind_Spots_of_Frontier_Models.fatima_blind_spot_challengeGot it. From now on I'll write everything inside Markdown blocks so you can copy easily.
Here is your full content entirely in Markdown:
# Fatima Fellowship 2026: Technical Challenge - Model Blind Spots
## 1. Model Overview
- **Model Tested:** [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Base Model)
- **Parameters:** 0.6B
- **Type:** Causal Language Model (Base / Pre-trained)
---
## 2. Methodology & Loading
To evaluate the model, I used **Google Colab** with a **T4… See the full description on the dataset page: https://huggingface.co/datasets/moseleydev/fatima_blind_spot_challenge.qwen3.5-0.8b-base-blind-spots
Qwen3.5-0.8B-Base Blind Spots Dataset
A curated dataset of 10 diverse failure cases from the Qwen/Qwen3.5-0.8B-Base model -- a 0.8B-parameter multimodal vision-language base model released in February 2026 by the Qwen Team.
Purpose
This dataset documents specific "blind spots" where the model produces incorrect, incoherent, or hallucinated outputs. Each example includes the input prompt, the expected correct output, the model's actual output, and a human-written analysis… See the full description on the dataset page: https://huggingface.co/datasets/Sindijow/qwen3.5-0.8b-base-blind-spots.qwen2.5-3b-blind-spots
Qwen2.5-3B Factual Recall Blind Spots
Model Tested
Qwen/Qwen2.5-3B
A 3.09B parameter base causal language model, pretrained only.
How I Loaded the Model
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_name = "Qwen/Qwen2.5-3B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.float16,
device_map="auto"
)
Loaded on Google Colab.… See the full description on the dataset page: https://huggingface.co/datasets/suhanii23/qwen2.5-3b-blind-spots.qwen3.5-2b-base-blind-spots
Qwen3.5-2B-Base Blind Spots Dataset
Model Tested
Qwen/Qwen3.5-2B-BaseReleased: March 2, 2026 | 2B parameters | Hybrid Gated DeltaNet + Attention architecture
Dataset Description
This dataset contains 12 input-output pairs where Qwen3.5-2B-Base makes incorrect or
unexpected predictions. Each row has:
category: the type of reasoning being tested
input: the prompt given to the model
expected_output: the correct answer
model_output: what the model actually… See the full description on the dataset page: https://huggingface.co/datasets/selssabilkadid/qwen3.5-2b-base-blind-spots.dmind-3-mini-blind-spots
DMind-3-mini Blind Spot Dataset
Model Tested
DMindAI/DMind-3-mini
Architecture: Qwen3.5-based
Parameters: 4B
Type: Domain-specific fine-tuned model (Web3/DeFi/Finance)
Primary Use: Computational Financial Actuary for DeFi analytics
Training: Fine-tuned on 82,000 high-value private financial samples
How I Loaded the Model
!pip install --upgrade transformers -q
!pip install torch accelerate -q
from transformers import AutoModelForCausalLM, AutoTokenizer… See the full description on the dataset page: https://huggingface.co/datasets/JackRabbit1122/dmind-3-mini-blind-spots.spotless-customer-service-training
Spotless Bin Co Customer Service Training Data
Training data for a customer service AI model for Spotless Bin Co, a residential trash can cleaning service.
Dataset Description
This dataset contains 8,776 conversational examples across 5 categories:
Category
Count
Description
FAQs
1,951
Frequently asked questions
Service
1,925
Service explanation dialogues
Objections
1,925
Objection handling examples
Booking
1,975
Booking flow conversations
Brand
1,000… See the full description on the dataset page: https://huggingface.co/datasets/rileyseaburg/spotless-customer-service-training.smollm2-blind-spots
SmolLM2-1.7B Blind Spots Dataset
This dataset contains 10 diverse examples where the SmolLM2-1.7B base model makes incorrect predictions or demonstrates "blind spots".
Model Tested
Model: SmolLM2-1.7B
Parameters: 1.7 Billion
Type: Base (Pre-trained)
How to Load the Model
The model was loaded using the transformers library in Python.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "HuggingFaceTB/SmolLM2-1.7B"
tokenizer =… See the full description on the dataset page: https://huggingface.co/datasets/Alhibb/smollm2-blind-spots.qwen3-base-blind-spots
Qwen3-0.6B-Base Blind Spots Dataset
Model tested: Qwen/Qwen3-0.6B-BaseParameters: 0.6B | Released: May 2025 | Type: Base (pretrained, not instruction-tuned)Tested by: Tito Osadebey | Platform: Google Colab (T4 GPU, free tier)
Overview
This dataset documents 10 diverse failure cases ("blind spots") identified in Qwen3-0.6B-Base through structured prompt testing. Failures span five categories: African geography and culture, temporal reasoning, arithmetic, logical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/titoausten/qwen3-base-blind-spots.qwen3-1.7b-blind-spots
Blind Spots of a Frontier Base Model: Failure Analysis of Qwen3-1.7B-Base
Dataset Link
Public Hugging Face Dataset
1. Model Selection
To complete this challenge, I browsed recently released open models on
Hugging Face within the 0.6B--6B parameter range.
The model selected for analysis:
Model Name: Qwen3-1.7B-Base Parameter Size: 1.7B Modality: Text (Causal Language Model) Type: Base model (not instruction-tuned) Availability: Public on Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Habibgm/qwen3-1.7b-blind-spots.qwen3-4b-base-blind-spots
Qwen3-4B-Base — Blind Spots Dataset
Model Tested
Model: Qwen/Qwen3-4B-Base
Type: Base model (pretraining only — no instruction tuning)
Parameters: 4.0B
Released: April 29, 2025
How the Model Was Loaded
Loaded on Google Colab (T4 GPU, free tier) using HuggingFace Transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B-Base")
model = AutoModelForCausalLM.from_pretrained(… See the full description on the dataset page: https://huggingface.co/datasets/saifeldein0/qwen3-4b-base-blind-spots.
