datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orthodox-patristic-corpus
Orthodox Patristic Corpus
Released on the Feast of the Triumph of Orthodoxy, First Sunday of Great Lent, 2026.
Dataset Summary
The Orthodox Patristic Corpus is a 116M-token pre-training corpus of Orthodox Christian
theological literature, assembled from the writings of 123 Church Fathers and Orthodox theologians
spanning the 1st through 20th centuries. The corpus is primarily in Russian, drawing on the
Azbyka.ru Orthodox digital library and other public-domain sources… See the full description on the dataset page: https://huggingface.co/datasets/jayfurzy/orthodox-patristic-corpus.Wolof-Non-Standard-Orthography
Dataset Description
Dataset Summary
This dataset contains pairs of non-standard and standard Wolof text, designed for training models to normalize informal Wolof writing found on social media, messaging apps, and online platforms.
The non-standard versions simulate real-world informal Wolof text with French code-switching, phonetic spellings, missing diacritics, and common typing variations.
The original Standard Wolof and English sentences are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Wolof-Non-Standard-Orthography.Orthography-and-Spelling
🇰🇿 Kazakh Orthographic and Quantifier Refinement
📖 Overview
This dataset contains 1,504 samples focusing on common orthographic and grammatical nuances in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
1,504
Total Words (approx.)
68,469
Avg. Words per Sample
45
Word Count Distribution (Per Field)
The following table details the distribution of word counts… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Orthography-and-Spelling.
