fatima-fellowship
Fatima-Fellowship-challenge
Dataset: Blind Spots of Frontier Models (Gemma-2-2b)
This dataset contains 12 diverse stress-test data points identifying the logical, physical, and ethical "blind spots" of the Gemma-2-2b base model. By using leading completions, we uncover how the model's raw weights handle reasoning, bias, and instruction drift without the safety layers of instruction tuning.
1. Model Tested
Model Name: google/gemma-2-2b
Parameters: 2.5 Billion
Type: Base Model (Causal Language Model)… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/Fatima-Fellowship-challenge.fatima-fellowship-blindspots
Dataset Card: Tiny-Aya-Base Amharic Evaluation Blindspots
Dataset Description
This dataset provides a targeted, interpretable checklist of reasoning failures and blindspots discovered in the CohereLabs/tiny-aya-base model when evaluated on Amharic language tasks across arithmetic, logic, science, history, and geography domains.
As models scale, evaluating their cross-lingual reasoning capabilities requires moving beyond aggregate metrics. This repository adopts a… See the full description on the dataset page: https://huggingface.co/datasets/MYGBM/fatima-fellowship-blindspots.fatima-fellowship-naruto-blip-captionsFatima_Fellowship
Bengali Quote Error Analysis Dataset
This repository contains an error-analysis dataset for Bengali quote understanding.
Model outputs: Fatima_Fellowship.csv
The goal is to document diverse model mistakes and propose a fine-tuning direction.
Model Tested
Model: Qwen/Qwen3.5-0.8B
Framework: transformers
Prompt format: chat template with a system instruction and one few-shot example
How the Model Was Loaded
The following code was used in the notebook:
from… See the full description on the dataset page: https://huggingface.co/datasets/SumitKumarKar01/Fatima_Fellowship.qwen-3-5-blindspots-fatima-fellowship
Qwen3.5-4B-Base Blindspots Dataset
A collection of 11 prompts where Qwen/Qwen3.5-4B-Base produces incorrect, incomplete, or degenerate outputs. Each row records the input prompt, the raw model output, the extracted model response, the expected (correct) response, and a label for the failure category.
Dataset Summary
Field
Value
Model tested
Qwen/Qwen3.5-4B-Base
Number of examples
11
Columns
prompt, raw_output, thinking, output, expected_output, blindspot… See the full description on the dataset page: https://huggingface.co/datasets/dawaawawa/qwen-3-5-blindspots-fatima-fellowship.fatima-fellowship-challenge-blind-spots-of-ministral-3-3b-base
Blind Spots of mistralai/Ministral-3-3B-Base-2512
Fatima Fellowship, Technical Challenge Submission
1. Model Selection
Model: mistralai/Ministral-3-3B-Base-2512
Parameters: ~3.8B (3.4B language model + 0.4B vision encoder)
Type: Base pre-trained, explicitly NOT fine-tuned for instructions or chat
Released: December 2025 | License: Apache 2.0
This model was selected because its model card explicitly states it is "the base pre-trained version, not fine-tuned for… See the full description on the dataset page: https://huggingface.co/datasets/Faiyaj/fatima-fellowship-challenge-blind-spots-of-ministral-3-3b-base.
