datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3.5-2b-base-blind-spots
Qwen3.5-2B-Base — Blind Spot Analysis (Text + Vision)
Model Tested
Field
Value
Model
Qwen/Qwen3.5-2B-Base
Parameters
2.27 B (2,274 M per HF metadata)
Architecture
Hybrid Gated-DeltaNet (dense FFN) — 24 LM layers (18 DeltaNet + 6 full-attention), ViT vision encoder
Type
Pre-trained base model (not instruction-tuned)
Context
262 144 tokens
Modalities
Text + Vision (early-fusion multimodal)
Key Contributions
Only multimodal… See the full description on the dataset page: https://huggingface.co/datasets/F555/qwen3.5-2b-base-blind-spots.medgemma-4b-hematologic-oncology-blind-spots
MedGemma Blind Spots: Hematologic Oncology & CAR-T Immunotherapy
A 13-probe red-team evaluation showing how Google's MedGemma-4B confidently hallucinates clinical-trial statistics, fabricates non-existent treatment regimens, and misdiagnoses lymphoma in hematologic oncology — a clinical domain absent from its documented training data.
Summary
This dataset documents failures of Google's MedGemma-4B on hematologic oncology prompts — a clinical subspecialty absent… See the full description on the dataset page: https://huggingface.co/datasets/Mateenah/medgemma-4b-hematologic-oncology-blind-spots.qwen3-base-blind-spots
Qwen3-0.6B-Base Blind Spots Dataset
Model tested: Qwen/Qwen3-0.6B-BaseParameters: 0.6B | Released: May 2025 | Type: Base (pretrained, not instruction-tuned)Tested by: Tito Osadebey | Platform: Google Colab (T4 GPU, free tier)
Overview
This dataset documents 10 diverse failure cases ("blind spots") identified in Qwen3-0.6B-Base through structured prompt testing. Failures span five categories: African geography and culture, temporal reasoning, arithmetic, logical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/titoausten/qwen3-base-blind-spots.qwen35-08b-base-blind-spots
Qwen3.5-0.8B-Base: Multi-Dimensional Blind Spot Dataset
Abstract
We present a structured dataset of blind spots discovered in
Qwen/Qwen3.5-0.8B-Base,
a 0.8B parameter base language model released in March 2026. Using an
automated pipeline grounded in three lines of NLP research — the
reversal curse (Berglund et al., ICLR 2024), confidence
calibration (Xiong et al., ICLR 2024), and behavioral testing
(Ribeiro et al., ACL 2020) — we probed 45 facts across 186
total prompts… See the full description on the dataset page: https://huggingface.co/datasets/Emmaka/qwen35-08b-base-blind-spots.tiny-aya-base-blind-spots
Blind Spots of a Frontier Base Model: Evaluation Dataset
This dataset documents blind spots discovered in a frontier open-weight base model through 19 structured evaluation tests. It was assembled as part of an assignment on identifying model weaknesses using the HelloBench evaluation framework.
Model Tested
CohereLabs/tiny-aya-base
Architecture: Transformer with Sliding Window Attention (SWA) (window size 4096, with RoPE) on three layers + one global attention layer… See the full description on the dataset page: https://huggingface.co/datasets/Mawube/tiny-aya-base-blind-spots.
