datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fattah-golden-superset
Fattah Golden
Fattah Golden is a large-scale, model-agnostic supervised fine-tuning (SFT) superset built by Nomeda Labs to train the Fattah family of coding and agentic coding models.
The dataset is designed as a labeled superset with no baked-in training ratios. This means the stored dataset is the complete cleaned and annotated corpus. Researchers and practitioners choose their own mixture at training time by filtering on the boolean capability columns.
Stats… See the full description on the dataset page: https://huggingface.co/datasets/nomeda-lab/fattah-golden-superset.old-nogay-turkish-ocr-corpus
Old Nogay Turkish OCR Corpus
This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis.
We are proud to release this as a contribution to Turkish NLP, Turkic-language NLP, historical NLP, and Turkish studies. The goal is not to pretend that a small OCR corpus is a polished benchmark. The goal is more… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/old-nogay-turkish-ocr-corpus.early-church-fathers
Early Church Fathers — Scripture Citation Index
68,240 passages from 349 Church Fathers, each keyed to the Bible verse it
comments on. Drawn from 20,253 distinct works and covering all 66 books.
This is a patristic catena in machine-readable form: given a verse, it returns
what the Fathers said about it. Nothing comparable exists as an open dataset —
the underlying translations are freely available, but the verse-level alignment
is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.fatwa-qa-evaluation
Fatwa QA Evaluation Dataset
Dataset Description
This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers.
Dataset Statistics
Total Samples: 2,000
Average Question Length: 243.9 characters
Average Answer Length: 492.3 characters
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.anadolu-ocr-corpus
Anadolu OCR Corpus
Anadolu OCR Corpus is an OpenCR export of OCR text and document metadata for 52 historical Ottoman Turkish, Turkish, and Arabic-containing PDF sources. The dataset is provided in two Hugging Face configs:
pages: one row per source page, including page-level OCR text, metadata, validation status, script direction, language detection, hashes, and split labels.
documents: one row per source document, including document-level concatenated text, markdown, aggregate… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/anadolu-ocr-corpus.fatwa-training_standardized_new
Fatwa Training Dataset (Standardized)
Dataset Description
This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs in a standardized conversation format for training Arabic language models. Each original sample has been augmented with 3 different prompt templates to increase training diversity.
Dataset Statistics
Total Samples: 9,953
Unique Fatwas: 6,212
Prompt Variations: 3 per fatwa
Average Question Length: 230.0… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-training_standardized_new.evliya-celebi-seyahatname-ocr
Evliya Celebi Seyahatname OCR Corpus
OCR-derived text from seven volumes of Evliya Celebi's Seyahatname, packaged
for corpus exploration, language modeling, OCR-quality analysis, and historical
Ottoman Turkish / Turkish NLP work.
Configs
pages: one row per OCR page, with page numbers and OCR status.
documents: one row per available volume, with page text concatenated.
Coverage
Available books: 1, 3, 4, 6, 7, 9, 10.
Missing from the 1-10 sequence: 2, 5, 8.… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/evliya-celebi-seyahatname-ocr.
