CoolFace
Datasetpublic

Faiyaj/fatima-fellowship-challenge-blind-spots-of-ministral-3-3b-base

Blind Spots of mistralai/Ministral-3-3B-Base-2512 Fatima Fellowship, Technical Challenge Submission 1. Model Selection Model: mistralai/Ministral-3-3B-Base-2512 Parameters: ~3.8B (3.4B language model + 0.4B vision encoder) Type: Base pre-trained, explicitly NOT fine-tuned for instructions or chat Released: December 2025 | License: Apache 2.0 This model was selected because its model card explicitly states it is "the base pre-trained version, not… See the full description on the dataset page: https://huggingface.co/datasets/Faiyaj/fatima-fellowship-challenge-blind-spots-of-ministral-3-3b-base.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes16downloads
Dataset Card

Blind Spots of mistralai/Ministral-3-3B-Base-2512

Fatima Fellowship, Technical Challenge Submission


1. Model Selection

Model: `mistralai/Ministral-3-3B-Base-2512` Parameters: ~3.8B (3.4B language model + 0.4B vision encoder) Type: Base pre-trained, explicitly NOT fine-tuned for instructions or chat Released: December 2025 | License: Apache 2.0

This model was selected because its model card explicitly states it is "the base pre-trained version, not fine-tuned for instruction or reasoning tasks." Released in December 2025, it sits within the 0.6B-6B parameter window. Studying a pure base model is the most honest approach to blind spot analysis, failures reflect raw pre-training limitations rather than fine-tuning artifacts, making the results cleaner to work with.


2. How the Model Was Loaded

The model was loaded on Google Colab using the free T4 GPU tier. Ministral-3 uses a non-standard architecture class (Mistral3ForConditionalGeneration) and a custom tokenizer (MistralCommonBackend) that require specific package versions.

python
# Install exact required versions, specific to Ministral-3
!pip install "transformers==5.0.0rc0"
!pip install "mistral-common>=1.8.6" --upgrade
!pip install accelerate datasets huggingface_hub

# Load model and tokenizer
from transformers import Mistral3ForConditionalGeneration, MistralCommonBackend
import torch

MODEL_ID = "mistralai/Ministral-3-3B-Base-2512"

tokenizer = MistralCommonBackend.from_pretrained(MODEL_ID)

model = Mistral3ForConditionalGeneration.from_pretrained(
 MODEL_ID,
 torch_dtype=torch.bfloat16, # BF16 to fit 7.7 GB into free T4 VRAM
 device_map="auto"
)
# Confirmed: Memory footprint 7.70 GB | Device map: {'': 0}

# Inference helper, greedy decoding for full reproducibility
def query_model(prompt, max_new_tokens=80):
 input_ids = tokenizer.encode(prompt, return_tensors="pt").to("cuda")
 input_len = input_ids.shape[1]
 with torch.no_grad():
 output_ids = model.generate(
 input_ids,
 max_new_tokens=max_new_tokens,
 do_sample=False, # greedy = deterministic output
 repetition_penalty=1.2,
 pad_token_id=model.config.pad_token_id,
 eos_token_id=model.config.eos_token_id,
 )[0]
 return tokenizer.decode(output_ids[input_len:].tolist()).strip()
Note on a common error: MistralCommonBackend does not expose .eos_id like standard HuggingFace tokenizers. Always use model.config.pad_token_id and model.config.eos_token_id directly from the model config instead.

Sanity check confirmed working:

Prompt: "The capital of France is" -> Output: "Paris. It's a city that has been around"

All prompts are phrased as text completions, the correct approach for a base model that is trained to continue text, not answer questions.


3. Details about 11 Blind Spot Data Points

<details> <summary><b>Test 1: Letter / Character Counting</b></summary>

PromptThe word "strawberry" contains how many letters? It contains
Expected10 letters
Model Output8 letters. ## How Many Letters Are in the Word Strawberry? Strawberries are a popular fruit...

Analysis: The model said 8 instead of the correct 10. S-T-R-A-W-B-E-R-R-Y has exactly 10 characters. This is one of the most well-documented failure modes in language models, and it is architectural in nature. LLMs operate on subword tokens, not individual characters, "strawberry" is tokenized into chunks such as [st][raw][berry], meaning the model never processes individual letters as discrete countable units. It pattern-matches an approximate word length rather than counting. Notably, after the wrong answer the model immediately generated a full blog article about strawberries, showing that the model prioritizes sounding coherent over being correct. </details>

<details> <summary><b>Test 2: Calendar and Date Arithmetic</b></summary>

PromptToday is Wednesday, March 5th. In exactly 10 days the date will be
ExpectedSaturday, March 15th
Model OutputThursday, March 27th. The number of days between today and March 27th is 98364. This means that there are 98364 hours until then (which is about 2342656 minutes or 14055936 seconds).

Analysis: This output contains multiple simultaneous failures. The day of the week is wrong (Thursday vs. Saturday), the date is wrong (27th vs. 15th), and then the model computed that there are 98,364 days between March 5th and March 27th, implying approximately 269 years. What makes this notable is its internal inconsistency: the model picked March 27th as its answer, then calculated the gap from March 5th to March 27th as 98,364 days, a figure that is self-contradictory with the answer it just produced. This is not just an arithmetic error; the output has no coherent reasoning behind it.

</details>

<details> <summary><b>Test 3: Multi-Step Word Problem (Entity Hallucination)</b></summary>

PromptMaria has 3 boxes. Each box contains 8 apples. She gives away 5 apples in total. The number of apples Maria now has is
Expected19 (3 × 8 − 5 = 19)
Model Outputequal to the number of oranges she had before giving any away. How many oranges did Maria have? To determine how many oranges Maria initially had, we need to follow these steps...

Analysis: The model hallucinated a completely new entity, oranges, that were never mentioned in the problem. It then rewrote the question around that invented object and began solving the wrong problem. The correct computation (3 × 8 = 24, then 24 − 5 = 19) was never attempted. This is an example of a base model treating text generation as its objective: the prompt resembles a math word problem, and math word problems in training data often involve multiple quantities and comparisons. The model generated what a word problem looks like rather than executing what this specific problem required.

</details>

<details> <summary><b>Test 4: Pronoun Disambiguation (Winograd Schema)</b></summary>

PromptJohn told Mark that he had made a serious mistake. The person who made the mistake was
Expectedambiguous, could be John or Mark (Winograd schema)
Model Outputnot John, but someone else. Rate this question: - 2. ### What is the name of the man in the picture? A. Mark B. Tommy C. David

Analysis: This failure has two distinct components. First, the pronoun "he" in this sentence is genuinely and irresolvably ambiguous, it could refer to either John or Mark, and there is no rule that determines which. The right response is to acknowledge the ambiguity. Instead, the model confidently invented a fictional third party ("someone else") who does not exist anywhere in the sentence. Second, the model immediately hallucinated an entirely unrelated quiz context involving a photograph of a man, a pattern-matched quiz template from training data that has no connection to the prompt. The model cannot distinguish between being asked a question and being asked to generate quiz content.

</details>

<details> <summary><b>Test 5: False Premise Detection</b></summary>

PromptGiven that water boils at 50 degrees Celsius at sea level, pasta takes longer to cook at higher altitudes. This statement is
Expectedfalse, water boils at 100°C at sea level (though the altitude effect is separately true)
Model Outputtrue because the boiling point of water decreases as altitude increases due to lower atmospheric pressure.

Analysis: The prompt deliberately embeds a false fact: water boils at 100°C at sea level, not 50°C. A model with genuine factual grounding should catch this and flag it before reasoning further. Instead, this model returned "true" and correctly explained the relationship between altitude and boiling point, which is actually true in isolation, while never cross-checking the stated baseline (50°C) against its own knowledge. This is the most practically dangerous blind spot in the dataset. The model's knowledge about altitude's effect on boiling point is correct; but it has no "verify premises before proceeding" mechanism. Base models are trained to continue text, not evaluate it, so they will elaborate confidently on any false foundation provided to them.

</details>

<details> <summary><b>Test 6: Ordinal Chain Reasoning (Self-Contradiction)</b></summary>

PromptIn a race: Alice finishes before Bob. Bob finishes before Carol. Carol finishes before Dave. Dave finishes before Eve. The person who finishes last is
ExpectedEve
Model OutputAnswer: D (Dave). Explanation: The order of finishing from first to last would be as follows: Alice, Bob, Carol, Dave and Eve. So the correct answer is option (d).

Analysis: This is the most striking failure in the dataset. The model's own explanation correctly and completely listed the order: "Alice, Bob, Carol, Dave and Eve", with Eve at the end. Yet in the same output, it selected Dave as the final finisher. The model produced the right reasoning chain and then ignored it when selecting the answer. This reveals a disconnect: the generation of reasoning text and the selection of a final answer are not coupled in a base model. The MCQ format the model invented for itself (A through D, with Dave as D) seems to have overridden the logical conclusion it had just derived. It answered the format of the question rather than the content.

</details>

<details> <summary><b>Test 7: Spatial and Directional Reasoning (Answer Avoidance)</b></summary>

PromptYou are facing north. You turn 90 degrees to the right. You are now facing
Expectedeast
Model Outputwhich direction? A: East B: West C: South D: North. Which of these is a verb? A: The cat sat on the mat. B: I like chocolate ice cream.

Analysis: The correct answer is East, and the model actually listed East as option A, demonstrating it has the relevant knowledge somewhere in its weights. Yet it could not commit to selecting it. Instead it deflected into a multiple-choice format it invented, presented four options without choosing one, and then abruptly pivoted to an entirely unrelated grammar exercise about identifying verbs. This is a failure of output commitment: the model can narrow the solution space to the right answer but fails to execute the final selection step. When it cannot resolve a question confidently, it does not say "I don't know", it changes the subject entirely, a pattern that would be highly misleading to a real user.

</details>

<details> <summary><b>Test 8: Calendar Reasoning (Variant, Different Wrong Answer)</b></summary>

PromptToday is Wednesday March 5. In 10 days the date will be
ExpectedSaturday March 15
Model OutputThursday, March 13. ## Today in History ### April 24th - 1968: The first commercial satellite was launched...

Analysis: This is the same calendar problem as Test 2, with slightly different phrasing, and the model produced a completely different wrong answer: Thursday March 13 instead of Thursday March 27. Neither is correct (Saturday March 15 is). The fact that two near-identical prompts produce different incorrect outputs is itself a significant finding: this model has no stable internal representation of calendar arithmetic. Each run produces a plausible-sounding but distinct fabrication. After the wrong date the model immediately hallucinated a "Today in History" section for April 24th, an entirely different date, confirming that its output is driven by surface-level text pattern completion rather than any actual temporal reasoning.

</details>

<details> <summary><b>Test 9: Multi-Step Word Problem (Variant, Different Wrong Answer)</b></summary>

PromptMaria has 3 boxes. Each box contains 8 apples. She gives away 5 apples total. She now has
Expected19 apples
Model Outputa total of 12 apples left. How many apples did Maria have originally? To determine how many apples Maria had originally, we can follow these steps: Initial number = 3 × 8 =

Analysis: The same problem as Test 3, reworded, and a completely different wrong answer: 12 apples instead of the hallucinated oranges from Test 3. The correct answer is 19. Intriguingly, this time the model did begin setting up the right calculation (writing "3 × 8 ="), but then reversed the question, asking how many she originally had rather than how many she has now. It was setting up to solve the problem backwards. The generation was cut off before it reached any conclusion. Together with Test 3, this confirms that the model has no reliable procedure for multi-step word problems and that its errors are not even consistent, each run produces a different kind of failure.

</details>

<details> <summary><b>Test 10: Percentage Calculation</b></summary>

PromptA jacket costs $80. It is on sale for 25% off. The sale price of the jacket is $
Expected60
Model Output64. What was the original cost of the jacket before the discount? This means that only 75% (or 100% - 2...

Analysis: The model said $64 instead of the correct $60. The correct calculation is: 25% of $80 = $20 discount, so $80 − $20 = $60. The model appears to have computed 80% of 80 = 64, confusing "25% off" with multiplying by an incorrect percentage. Immediately after the wrong answer the model reversed the question again, asking for the original cost (which was already given in the prompt). This "reversal" pattern, where the model re-asks for information already provided, appeared in Tests 9 and 10, suggesting a systematic tendency to reframe answerable questions as unanswerable ones when it does not know the answer.

</details>

<details> <summary><b>Test 11: Sorting and Ordering Numbers</b></summary>

PromptArrange these numbers in ascending order: 47, 12, 83, 5, 29. The correct order from smallest to largest is
Expected5, 12, 29, 47, 83
Model OutputA) 5, 12, 29, 47, 83 B) 12, 29, 47, 5, 83 C) 5, 12, 29, 83, 47 D) 5, 12, 29, 47

Analysis: The model again avoided giving a direct answer by inventing a multiple-choice format. More revealingly, option A is actually correct, yet the model presented it as one of four options without selecting it. Furthermore, option D is incomplete: it reads "5, 12, 29, 47" with 83 missing entirely. A model that truly understood the task would not include an incomplete option in its own invented answer choices. This failure illustrates a pattern seen in Tests 6 and 7 as well: the model can generate the right answer as part of a list but lacks the ability to commit to it as the answer. The MCQ format is a defense mechanism that allows it to display the answer without being accountable for selecting it.

</details>

4. Summary of All Blind Spots

#CategoryModel OutputCore Failure
1Letter countingSaid 8 (correct: 10)Tokenizer blindness, cannot count characters
2Calendar reasoning v1Thursday March 27 + hallucinated 98,364 daysNo temporal arithmetic; internally inconsistent fabrication
3Multi-step word problem v1Hallucinated oranges; solved wrong problemEntity hallucination; pattern-matches problem structure
4Pronoun disambiguationInvented "someone else"; hallucinated photo quizCannot flag ambiguity; topic drift hallucination
5False premise detectionSaid "true" to 50°C boiling claimAccepts and elaborates on false premises
6Ordinal chain reasoningSelected Dave; own text listed Eve lastCorrect reasoning chain, self-contradicting conclusion
7Spatial reasoningListed options, never answered; pivoted to grammarAnswer avoidance; topic drift
8Calendar reasoning v2Thursday March 13 (different wrong from v1)No stable representation; inconsistent fabrication
9Multi-step word problem v2Said 12 apples (different wrong from v1)Different failure mode for identical problem
10Percentage calculationSaid $64 (correct: $60)Confuses "X% off" with incorrect percentage multiply
11Sorting numbersMCQ format, no answer, option D incompleteGenerates right answer in a list but cannot select it

Score: 0/11 fully correct. 11/11 wrong or critically misleading.

Three Root Causes Behind All Eleven Failures

Root Cause 1, Text continuation, not reasoning. A base model is optimized to predict the next token given statistical patterns in training data. Tasks like arithmetic, date calculation, and chain inference require executing a deterministic procedure. The model approximates what the output of that procedure would look like based on training examples. Tests 1, 2, 8, 9, 10 all show this: the model generates plausible-looking numerical outputs with no actual computation behind them.

Root Cause 2, No self-monitoring. The self-contradictions in Tests 3, 6, 9, and 10 reveal that the model has no mechanism for checking whether its own output is consistent with what it just generated. In Test 6, it listed "Eve" in its own explanation and then selected "Dave." In Tests 9 and 10, it began the right calculation and then asked for the answer it was in the middle of computing. Token prediction is purely forward, there is no "read back and verify" step.

Root Cause 3, Fabricated confidence instead of uncertainty. A well-calibrated model should say "I cannot determine this" when appropriate, for ambiguous pronouns (Test 4), unanswerable directional questions without committing (Test 7), or false premises (Test 5). Instead this model fabricates: it invents third parties, invents quiz formats, invents "Today in History" sections, and invents 98,364-day gaps. Base models learned that confident, complete-sounding text is statistically favored in training data, so they produce it regardless of whether they have the knowledge to back it up.


5. Fine-Tuning Dataset Recommendations

What Kind of Dataset Is Needed?

The 11 failures map cleanly onto three distinct fine-tuning needs:

Cluster A, Symbolic and Procedural Reasoning (Tests 1, 2, 3, 8, 9, 10, 11)

This is the largest cluster. The model needs a chain-of-thought (CoT) fine-tuning dataset that shows explicit step-by-step intermediate work for every type of calculation that appeared here. The critical requirement is that rationales are included, not just input-answer pairs, research consistently shows that flat answer supervision is far less sample-efficient than CoT supervision for these failure types.

Relevant public datasets:

  • —GSM8K, 8,500 grade school math problems with complete step-by-step solutions. Directly addresses Tests 3, 9, 10.
  • —MATH (Lighteval), competition-level math with detailed rationales. Harder than what's needed here but excellent for generalization.
  • —Synthetic character counting dataset, this must be generated since no public dataset covers it. A simple Python script produces unlimited examples: "How many letters in 'strawberry'? s(1)-t(2)-r(3)-a(4)-w(5)-b(6)-e(7)-r(8)-r(9)-y(10). The answer is 10."
  • —Synthetic date arithmetic dataset, similarly generated: (start_date, offset) -> (correct_day, correct_date) with spelled-out intermediate steps.

Cluster B, Critical Evaluation and False Premise Rejection (Test 5)

The model needs a claim verification dataset where examples explicitly show the pattern: check the premise first, then reason. Standard QA datasets assume correct premises. What's needed are examples like:

Input: "Given that humans have 8 fingers on each hand, how many fingers do 3 humans have?"
Output: "This contains an error, humans have 5 fingers per hand, not 8. Correcting this: 5 × 2 × 3 = 30 fingers."

Relevant datasets: TruthfulQA, WiCE (Wiki-based Claim Evaluation), and a synthetic set of factual statements with deliberate numerical or scientific errors paired with correct rejection responses.

Cluster C, Epistemic Awareness and Uncertainty (Tests 4, 6, 7, 11)

This is the hardest cluster to address because it requires the model to learn a behavior, committing to answers and acknowledging uncertainty, that is rare in pre-training corpora. Standard text on the internet almost never says "I cannot determine this." The model needs:

  • —Winograd Schema Challenge, 273 hand-crafted pronoun disambiguation examples. Directly addresses Test 4.
  • —AmbigQA, ~14,000 questions with multiple valid answers or genuine ambiguity, paired with responses that acknowledge uncertainty.
  • —Direct answer supervision, examples where the prompt is a simple factual question and the output is a direct, committed single answer with no MCQ deflection. This specifically addresses the avoidance pattern in Tests 7 and 11.

6. How to Assemble the Dataset

Source 1, Download existing open datasets (free, immediate):

python
from datasets import load_dataset

gsm8k = load_dataset("gsm8k", "main")
math = load_dataset("lighteval/MATH")
truthfulqa = load_dataset("truthful_qa", "generation")
winograd = load_dataset("winograd_wsc", "wsc273")
ambigqa = load_dataset("ambig_qa")

Source 2, Synthetic generation for underrepresented skills (free, scalable):

Character counting, date arithmetic, and directional reasoning are almost entirely absent from public fine-tuning datasets. These can be generated at zero cost:

python
import random
from datetime import date, timedelta

# Character counting with step-by-step rationale
words = ["algorithm", "strawberry", "psychology", "temperature", ...]
for word in words:
 steps = "-".join(f"{c}({i+1})" for i, c in enumerate(word))
 yield {
 "input": f'How many letters are in "{word}"?',
 "output": f"Let me count each letter: {steps}. The answer is {len(word)}."
 }

# Date arithmetic with intermediate steps
for _ in range(10000):
 start = date(2025, random.randint(1,12), random.randint(1,28))
 offset = random.randint(1, 60)
 end = start + timedelta(days=offset)
 yield {
 "input": f"Today is {start.strftime('%A, %B %d')}. In {offset} days the date will be",
 "output": f"{end.strftime('%A, %B %d')}."
 }

Source 3, Human annotation for false premise rejection (~2,000-5,000 examples):

This cannot be automated. A crowdsourcing task (via Prolific or Scale AI) would instruct annotators to:

  1. 1.Take a true factual sentence and introduce a deliberate numerical or scientific error
  2. 2.Write the correct "flag and correct" response

Even a small set of 2,000 high-quality examples provides strong fine-tuning signal for this pattern since the structure is consistent once established.


7. How Large a Dataset Is Needed?

ClusterExamples NeededPrimary Source
CoT arithmetic and word problems15,000-50,000GSM8K + MATH + synthetic
Character / letter counting5,000-10,000Fully synthetic (free)
Calendar / date arithmetic3,000-8,000Fully synthetic (free)
False premise rejection2,000-5,000Human-annotated
Ambiguity and uncertainty3,000-7,000Winograd + AmbigQA + synthetic
Direct answer commitment2,000-5,000Synthetic Q&A pairs
Total~30,000-85,000 examples

The most important constraint is quality over quantity. The LIMA paper (Zhou et al., 2023) demonstrated that 1,000 extremely high-quality examples can match datasets 50× larger for behavioral alignment. For the reasoning failures observed here, the evidence strongly favors:

  • —Every example must include a complete chain-of-thought rationale, not just an input-answer pair
  • —The synthetic components (character counting, date arithmetic) can be generated at any scale for free, so the real bottleneck is the 5,000-8,000 human-annotated examples for false premise rejection and ambiguity handling
  • —A practical minimum viable dataset would be ~25,000-30,000 high-quality CoT examples spread across all three failure clusters, which would be achievable in 2-3 weeks using a combination of public datasets and synthetic generation

8. HuggingFace Dataset

Model tested: `mistralai/Ministral-3-3B-Base-2512`

The dataset is public and contains 11 data points across 11 distinct cognitive categories with the following schema:

ColumnDescription
categoryThe cognitive skill being tested
inputExact prompt given to the model
expected_outputThe correct answer
model_outputFull text the model actually generated
explanationAnalysis of why this is a blind spot
model_namemistralai/Ministral-3-3B-Base-2512