NimsW/nanbeige-3b-blindspots-eval
Technical Challenge: Blind Spots of Frontier Models Author: Nimshi Wanniarachchi This dataset documents systematic failure cases observed when evaluating a recent open-source base language model with approximately 0.6–6B parameters. The dataset contains 10 diverse evaluation examples, each including: Input prompt Expected output Actual model output Model tested Model Name: Nanbeige4-3B-Base Model Link: https://huggingface.co/Nanbeige/Nanbeige4-3B-Base… See the full description on the dataset page: https://huggingface.co/datasets/NimsW/nanbeige-3b-blindspots-eval.
This repository belongs to NimsW on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
