datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtu-llm-2b-base-blind-spots
Youtu-LLM-2B-Base Blind Spots Evaluation Dataset
This dataset contains 75 evaluation prompts used to analyze the failure modes of tencent/Youtu-LLM-2B-Base,
a 1.96B parameter dense base language model released on December 31, 2025. Each row includes the input prompt, the expected answer, and the model’s
generated output obtained during inference on a Google Colab T4 GPU.
The prompts span 13 broad categories including arithmetic, logic, multilingual generation, instruction following… See the full description on the dataset page: https://huggingface.co/datasets/k-imtz/youtu-llm-2b-base-blind-spots.youtu-llm-2b-base-blindspots
Blind Spots of tencent/Youtu-LLM-2B-Base
Model Tested
Model: tencent/Youtu-LLM-2B-BaseLink: https://huggingface.co/tencent/Youtu-LLM-2B-Base
This evaluation was conducted on the base (pretrained) version of the model, not an instruction-tuned variant.
Objective
The goal of this dataset is to identify systematic failure patterns ("blind spots") of the Youtu-LLM-2B-Base model through targeted probing. The evaluation focuses on arithmetic reasoning, unit… See the full description on the dataset page: https://huggingface.co/datasets/Corneille1/youtu-llm-2b-base-blindspots.youtu-llm-2b-base-blindspots
Youtu-LLM-2B-Base Blind Spots
This dataset contains 10 failure cases collected while probing tencent/Youtu-LLM-2B-Base, an open base language model on Hugging Face. The model card describes it as a Base release, lists it at 1.96B parameters, and notes support for 131,072 context length.
Model tested
Model: tencent/Youtu-LLM-2B-Base
Model type: Base model
Parameters: 1.96B
Context length: 131,072
I selected this model because it fit the assignment constraints well: it is… See the full description on the dataset page: https://huggingface.co/datasets/Candace352/youtu-llm-2b-base-blindspots.
