datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Blind_Spots_of_Frontier_Models
Qwen3-0.6B-Base — Blind Spots Dataset
Model Tested
Qwen/Qwen3-0.6B-Base
Type: Causal Language Model (base / pretraining only — not instruction-tuned)
Parameters: 0.6B (0.44B non-embedding)
Released: April–May 2025 by Alibaba Cloud's Qwen Team
Context Length: 32,768 tokens
How the Model Was Loaded
The model was loaded in a Google Colab T4 GPU notebook using HuggingFace transformers >= 4.51.0(required because the qwen3 architecture key was added in… See the full description on the dataset page: https://huggingface.co/datasets/Pidoxy/Blind_Spots_of_Frontier_Models.nigerian-food-science-frontier-blindspotsModel Tested
I evaluated SmolLM2-1.7B. This is a compact base model (not fine-tuned for instruction following), which allows for the identification of raw data biases and "blind spots" in its pre-training corpus regarding specialized regional knowledge.
How I Loaded and Tested the Model
The model was loaded in a Google Colab environment using a T4 GPU. I used the transformers library with bfloat16 precision to ensure efficient memory management while maintaining numerical accuracy.… See the full description on the dataset page: https://huggingface.co/datasets/Isaac101111/nigerian-food-science-frontier-blindspots.
