Youtu-LLM-2B
youtu-llm-2b-base-blind-spots
Youtu-LLM-2B-Base Blind Spots Evaluation Dataset
This dataset contains 75 evaluation prompts used to analyze the failure modes of tencent/Youtu-LLM-2B-Base,
a 1.96B parameter dense base language model released on December 31, 2025. Each row includes the input prompt, the expected answer, and the model’s
generated output obtained during inference on a Google Colab T4 GPU.
The prompts span 13 broad categories including arithmetic, logic, multilingual generation, instruction following… See the full description on the dataset page: https://huggingface.co/datasets/k-imtz/youtu-llm-2b-base-blind-spots.youtu-llm-2b-base-blind-spots
Youtu-LLM-2B-Base Blind Spots Dataset
What is this?
I tested a small AI language model called Youtu-LLM-2B-Base (made by Tencent) to find places where it gives wrong or strange answers. I gave it 50 different questions and kept the 25 cases where it clearly failed.
This dataset contains those 25 failures — the question I asked, what the correct answer should be, what the model actually said, and why it was wrong.
About the Model
Name: Youtu-LLM-2B-Base
Link:… See the full description on the dataset page: https://huggingface.co/datasets/Afras/youtu-llm-2b-base-blind-spots.youtu-llm-2b-base-blindspots
Blind Spots of tencent/Youtu-LLM-2B-Base
Model Tested
Model: tencent/Youtu-LLM-2B-BaseLink: https://huggingface.co/tencent/Youtu-LLM-2B-Base
This evaluation was conducted on the base (pretrained) version of the model, not an instruction-tuned variant.
Objective
The goal of this dataset is to identify systematic failure patterns ("blind spots") of the Youtu-LLM-2B-Base model through targeted probing. The evaluation focuses on arithmetic reasoning, unit… See the full description on the dataset page: https://huggingface.co/datasets/Corneille1/youtu-llm-2b-base-blindspots.youtu-llm-2b-base-blindspots
Youtu-LLM-2B-Base Blind Spots
This dataset contains 10 failure cases collected while probing tencent/Youtu-LLM-2B-Base, an open base language model on Hugging Face. The model card describes it as a Base release, lists it at 1.96B parameters, and notes support for 131,072 context length.
Model tested
Model: tencent/Youtu-LLM-2B-Base
Model type: Base model
Parameters: 1.96B
Context length: 131,072
I selected this model because it fit the assignment constraints well: it is… See the full description on the dataset page: https://huggingface.co/datasets/Candace352/youtu-llm-2b-base-blindspots.
