CoolFace
Datasetpublic

NimsW/nanbeige-3b-blindspots-eval

Technical Challenge: Blind Spots of Frontier Models Author: Nimshi Wanniarachchi This dataset documents systematic failure cases observed when evaluating a recent open-source base language model with approximately 0.6–6B parameters. The dataset contains 10 diverse evaluation examples, each including: Input prompt Expected output Actual model output Model tested Model Name: Nanbeige4-3B-Base Model Link: https://huggingface.co/Nanbeige/Nanbeige4-3B-Base… See the full description on the dataset page: https://huggingface.co/datasets/NimsW/nanbeige-3b-blindspots-eval.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
0likes3downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
NimsW/nanbeige-3b-blindspots-eval · CoolFace