datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fattah-golden-superset
Fattah Golden
Fattah Golden is a large-scale, model-agnostic supervised fine-tuning (SFT) superset built by Nomeda Labs to train the Fattah family of coding and agentic coding models.
The dataset is designed as a labeled superset with no baked-in training ratios. This means the stored dataset is the complete cleaned and annotated corpus. Researchers and practitioners choose their own mixture at training time by filtering on the boolean capability columns.
Stats… See the full description on the dataset page: https://huggingface.co/datasets/nomeda-lab/fattah-golden-superset.Archer-Code-1.5B
✨ ArcherCodeR
🏹️ Reinforcement Learning for Enhanced Code Reasoning in LLMs 🎯
Overview
ArcherCodeR-Dataset is a dataset of verifiable, challenging, and diverse coding questions (6.7K). This dataset is used to train the ArcherCodeR model series, which consists of code reasoning models trained using large-scale rule-based reinforcement learning with carefully designed datasets and training recipes.
We select, clean, and curate coding problems from… See the full description on the dataset page: https://huggingface.co/datasets/Fate-Zero/Archer-Code-1.5B.old-nogay-turkish-ocr-corpus
Old Nogay Turkish OCR Corpus
This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis.
We are proud to release this as a contribution to Turkish NLP, Turkic-language NLP, historical NLP, and Turkish studies. The goal is not to pretend that a small OCR corpus is a polished benchmark. The goal is more… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/old-nogay-turkish-ocr-corpus.ArcherCodeR-Dataset
✨ ArcherCodeR
🏹️ Reinforcement Learning for Enhanced Code Reasoning in LLMs 🎯
Overview
ArcherCodeR-Dataset is a dataset of verifiable, challenging, and diverse coding questions (6.7K). This dataset is used to train the ArcherCodeR model series, which consists of code reasoning models trained using large-scale rule-based reinforcement learning with carefully designed datasets and training recipes.
We select, clean, and curate coding problems from… See the full description on the dataset page: https://huggingface.co/datasets/Fate-Zero/ArcherCodeR-Dataset.early-church-fathers
Early Church Fathers — Scripture Citation Index
68,240 passages from 349 Church Fathers, each keyed to the Bible verse it
comments on. Drawn from 20,253 distinct works and covering all 66 books.
This is a patristic catena in machine-readable form: given a verse, it returns
what the Fathers said about it. Nothing comparable exists as an open dataset —
the underlying translations are freely available, but the verse-level alignment
is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.Fattah-Orchestrator-Dataset
Fattah Orchestrator Dataset
A supervised fine-tuning dataset for training LLMs to act as orchestrators inside AI coding agents. The model receives a coding request written in Egyptian Arabic and must produce a structured JSON plan: a brief reasoning trace, a request summary, and an ordered list of dependency-aware subtasks.
Purpose
Egyptian-Arabic-speaking developers often prompt coding agents in colloquial Egyptian Arabic (not Modern Standard Arabic). Off-the-shelf LLMs… See the full description on the dataset page: https://huggingface.co/datasets/nomeda-lab/Fattah-Orchestrator-Dataset.fatwa-qa-evaluation
Fatwa QA Evaluation Dataset
Dataset Description
This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers.
Dataset Statistics
Total Samples: 2,000
Average Question Length: 243.9 characters
Average Answer Length: 492.3 characters
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.Fathom-V0.4-RL-CompressionFathom-V0.4-SFT-Shortest-ChainsFathom-V0.6-Iterative-Curriculum-Learningblind-spot-fatima-institute-qwen3.5-0.8b
Evaluation Report of Qwen3.5-0.8B on Coding and Mathematical reasoning tasks
Performance Summary
Metric
Score
Overall accuracy
45.8%
Coding accuracy
33.3%
Math accuracy
58.3%
Total tests evaluated
24
Coding tests (12 total)
Result
Count
Correct
4
Partially correct
1
Incorrect
7
Math tests (12 total)
Result
Count
Correct
7
Partially correct
2
Incorrect
3
What the Model Did Well… See the full description on the dataset page: https://huggingface.co/datasets/Acesif/blind-spot-fatima-institute-qwen3.5-0.8b.anadolu-ocr-corpus
Anadolu OCR Corpus
Anadolu OCR Corpus is an OpenCR export of OCR text and document metadata for 52 historical Ottoman Turkish, Turkish, and Arabic-containing PDF sources. The dataset is provided in two Hugging Face configs:
pages: one row per source page, including page-level OCR text, metadata, validation status, script direction, language detection, hashes, and split labels.
documents: one row per source document, including document-level concatenated text, markdown, aggregate… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/anadolu-ocr-corpus.fatima-techchallenge-blindspots-2026
Overview
This dataset highlights blind spots and errors of the open-source language model Qwen3-0.6B, released on Hugging Face. It contains 10 diverse prompts across multiple categories:
Factual knowledge
Ambiguous or tricky questions
Commonsense reasoning
Hardware and robotics knowledge
Local and African context
Arithmetic
Local language translation (Kenyan Swahili/Slang)
Simple/edge cases
Code generation
Cultural or language understanding
Each entry includes the input prompt, the… See the full description on the dataset page: https://huggingface.co/datasets/eustacelee/fatima-techchallenge-blindspots-2026.arabic-ocr-jokes
Extracted Arabic Jokes Dataset
Dataset Description
This dataset contains Arabic jokes extracted and cleaned from various PDF booklets using OCR and LLM-based post-processing. It was created to support research in Arabic Humor Generation.
Data Fields
Joke_Text: The full text of the joke in Arabic.
Source
Extracted from public domain joke books.
smollm3-3b-base-blind-spots
SmolLM3-3B-Base Blind Spots Dataset
This dataset contains 10 test cases where I explored the failure modes of
SmolLM3-3B-Base,
a 3 billion parameter base language model released by HuggingFace in 2025.
The goal was to find diverse cases where the model makes clearly incorrect
or unexpected completions its "blind spots."
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
Parameters: 3B
Type: Base pretrained model
License: Apache 2.0
How I Loaded the Model
I… See the full description on the dataset page: https://huggingface.co/datasets/FatimaAfzal01/smollm3-3b-base-blind-spots.fatima_blind_spot_challengeGot it. From now on I'll write everything inside Markdown blocks so you can copy easily.
Here is your full content entirely in Markdown:
# Fatima Fellowship 2026: Technical Challenge - Model Blind Spots
## 1. Model Overview
- **Model Tested:** [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Base Model)
- **Parameters:** 0.6B
- **Type:** Causal Language Model (Base / Pre-trained)
---
## 2. Methodology & Loading
To evaluate the model, I used **Google Colab** with a **T4… See the full description on the dataset page: https://huggingface.co/datasets/moseleydev/fatima_blind_spot_challenge.MineSafety-QA-Dataset
矿山安全领域 QA 数据集
基于中国矿山安全法规构建的问答对数据集,用于 QLoRA 领域微调。
数据来源
《煤矿安全规程》(2025)
《金属非金属矿山安全规程》(2020)
数据规模
原始生成:7874 条
AI 质量评估过滤后:7265 条
数据格式
Alpaca 格式,包含 <think> 推理链:
{
"instruction": "问题",
"input": "",
"output": "<think>\n推理过程...\n</think>\n\n正式回答...",
"system": "你是一位精通中国矿山安全法律法规的资深专家..."
}
构建流程
PDF 规程文档经 MinerU 转为 Markdown
Easy Dataset 自动分块、提取问题、生成答案(DeepSeek-R1-0528-Qwen3-8B)
AI 自动评分(满分 5 分),过滤 3.5 分以下的低质量 QA 对… See the full description on the dataset page: https://huggingface.co/datasets/FateDefier/MineSafety-QA-Dataset.qwen35-2b-base-blindspots-fatima
Qwen3.5-2B-Base Blind Spots for Exact-Answer Reasoning
This dataset contains incorrect predictions made by Qwen/Qwen3.5-2B-Base on a diverse exact-answer benchmark created for the 2026 Fatima Fellowship technical challenge.
The benchmark is designed to expose blind spots in a recent base model under tasks where there is a single correct output and correctness can be measured exactly. The dataset stores only failure cases: each row includes the question, the expected answer, and the… See the full description on the dataset page: https://huggingface.co/datasets/DrNerd/qwen35-2b-base-blindspots-fatima.qwen-3-5-blindspots-fatima-fellowship
Qwen3.5-4B-Base Blindspots Dataset
A collection of 11 prompts where Qwen/Qwen3.5-4B-Base produces incorrect, incomplete, or degenerate outputs. Each row records the input prompt, the raw model output, the extracted model response, the expected (correct) response, and a label for the failure category.
Dataset Summary
Field
Value
Model tested
Qwen/Qwen3.5-4B-Base
Number of examples
11
Columns
prompt, raw_output, thinking, output, expected_output, blindspot… See the full description on the dataset page: https://huggingface.co/datasets/dawaawawa/qwen-3-5-blindspots-fatima-fellowship.Synanthic-Arabic-jokes-llama-3.3-70b-versatile
Dataset Card for Synanthic-Arabic-jokes-60k
Dataset Description
This dataset contains a large collection of synthetically generated Arabic jokes. The dataset was designed to capture a wide variety of Arabic dialects, topics, and comedic styles while maintaining strict content safety guidelines.
Data Generation
The jokes were generated using a programmatic loop that randomly combined a specific topic, dialect, and style into a prompt. The generation process… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Synanthic-Arabic-jokes-llama-3.3-70b-versatile.fatwa-training_standardized_new
Fatwa Training Dataset (Standardized)
Dataset Description
This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs in a standardized conversation format for training Arabic language models. Each original sample has been augmented with 3 different prompt templates to increase training diversity.
Dataset Statistics
Total Samples: 9,953
Unique Fatwas: 6,212
Prompt Variations: 3 per fatwa
Average Question Length: 230.0… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-training_standardized_new.Fatimah_Fellowship_Blind_Spot
Qwen3-4B-Base Blind Spots Dataset
Model Tested
Model: Qwen/Qwen3-4B-Base
Type: Causal Language Model — pretrained base model (NOT instruction-tuned)
Parameters: 4.0 billion (3.6B non-embedding)
Architecture: 36 layers, 32 attention heads (GQA: 32 Q / 8 KV)
Context Length: 32,768 tokens
Training: 36 trillion tokens across 119 languages in a 3-stage pretraining pipeline
Overview
This dataset documents 10 confirmed blind spots of Qwen3-4B-Base identified… See the full description on the dataset page: https://huggingface.co/datasets/abdulmatinomotoso/Fatimah_Fellowship_Blind_Spot.evliya-celebi-seyahatname-ocr
Evliya Celebi Seyahatname OCR Corpus
OCR-derived text from seven volumes of Evliya Celebi's Seyahatname, packaged
for corpus exploration, language modeling, OCR-quality analysis, and historical
Ottoman Turkish / Turkish NLP work.
Configs
pages: one row per OCR page, with page numbers and OCR status.
documents: one row per available volume, with page text concatenated.
Coverage
Available books: 1, 3, 4, 6, 7, 9, 10.
Missing from the 1-10 sequence: 2, 5, 8.… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/evliya-celebi-seyahatname-ocr.fatima_institute_blind_spot
Blind Spots of Nanbeige/Nanbeige4-3B-Base
1. Model Tested
Nanbeige/Nanbeige4-3B-Base
Field
Detail
Released
December 13, 2025
Parameters
~3B
Type
TRUE BASE MODEL — pre-trained only on 23 trillion tokens, no SFT, no RLHF
Languages
English + Chinese (primary), multilingual coverage
License
Apache 2.0
2. How the Model Was Loaded
The model was loaded on Google Colab (free tier, T4 GPU, 16 GB VRAM) using the Hugging Face transformers… See the full description on the dataset page: https://huggingface.co/datasets/Nabeelah04/fatima_institute_blind_spot.fatimah-falcon3-3b-base-blindspots
Falcon3-3B-Base — Blind Spot Evaluation Dataset
This dataset documents 10 confirmed, diverse failure modes ("blind spots") of
tiiuae/Falcon3-3B-Base, a raw
pretrained 3B causal language model released by the Technology Innovation Institute
in December 2024. Every data point was empirically verified — the model was run live
and confirmed to produce the wrong output before inclusion.
Model Under Test
Field
Value
Model
tiiuae/Falcon3-3B-Base
Size
3B… See the full description on the dataset page: https://huggingface.co/datasets/Abdulsobur/fatimah-falcon3-3b-base-blindspots.icaire-ai-glossary-enriched
ICAIRE AI Glossary — Enriched (Mustalih Living)
Bilingual Arabic-English AI glossary based on the ICAIRE canonical vocabulary,
enriched through a multi-layer LLM pipeline into a fully structured
multimodal dataset: metaphors, detailed explanations, UML diagrams, typed
knowledge-graph edges, and narrator-voice story-track assignments.
Dataset structure
Each term (1,242 total) is one JSON record with these fields:
Field
Type
Description
english_term
string… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/icaire-ai-glossary-enriched.fatima-submission-dataset
Fatima Fellowship - Blind Spots Dataset
This dataset is used to store some blind spots on the base model for fellowship submission 2026
Model Tested
CohereLabs/tiny-aya-base
The model was tested primarily on Indonesian language queries to evaluate its capabilities and limitations on low-resource Southeast Asian language content.
Each data point contains the input prompt, the expected output, the model's actual output, and a classification of the error type.
How… See the full description on the dataset page: https://huggingface.co/datasets/habibi26/fatima-submission-dataset.
