datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenITI-MAKHZAN-Ottoman-Lines
Dataset Card for OpenITI MAKHZAN Ottoman Lines
Dataset Summary
This dataset contains line-level image-text pairs of historical Ottoman Turkish manuscripts and printed documents. It is derived from the OpenITI MAKHZAN dataset, a large aggregation of Arabic-script ground truth and evaluation data developed by the Open Islamicate Texts Initiative (OpenITI).
The dataset specifically focuses on Ottoman Turkish texts and is highly valuable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OpenITI-MAKHZAN-Ottoman-Lines.Akis-Ottoman-Dataset
Akis-Dataset
This repository provides the test dataset used in the paper "Automatic Transcription of Ottoman Documents Using Deep Learning". It contains line segment images of Ottoman documents along with their corresponding transcriptions.
Dataset Overview
The dataset contains 8,037 image–transcription pairs of Ottoman handwritten document line segments.
Format 1: HuggingFace Dataset (Parquet — Recommended)
The dataset is natively available as a… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/Akis-Ottoman-Dataset.ottoman-place-names-gazetteer
Ottoman Turkish Place Names Gazetteer (Transliteration Dataset)
Dataset Summary
This dataset serves as a specialized parallel corpus for Ottoman Turkish to Modern Turkish Latin script transliteration, focusing specifically on historical place names (toponyms). It is designed to enhance the performance of Large Language Models (LLMs) and OCR post-processing tools in recognizing and correctly transcribing historical geographical entities.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/ottoman-place-names-gazetteer.OsmanlicaEkmekveNisastaKitabi
Osmanlıca Ekmek ve Nişasta Kitabı Veri Seti
Bu veri seti, 1331 (1915) yılında Matbaa-i Âmire tarafından basılan, İzmir mebusu ve Dârülmuallimât sanayi-i ziraiye muallimi İhsan Bey tarafından yazılan "Kadınlara Amelî Sanayi-i Ziraiye Dersleri: Cüz 5 - Ekmek ve Nişastacılık Sanatı" kitabının görsellerini, Osmanlıca transkripsiyonunu ve (parantez içinde) günümüz Türkçesi sadeleştirmelerini/çevirilerini içerir.
Veri Seti Yapısı
image: Kitap sayfasının orijinal yüksek… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OsmanlicaEkmekveNisastaKitabi.rlvrCHURRO-Ottoman-Turkish-Subset
CHURRO Ottoman Turkish Subset
This repository contains the Ottoman Turkish historical document subset extracted from the CHURRO-DS dataset published at EMNLP 2025.
CHURRO is a 3B-parameter open-weight Large Vision-Language Model (VLM) specialized for high-accuracy, low-cost historical text recognition across diverse scripts and historical variants.
📊 Dataset Overview
Total Samples: 237 historical manuscript page images with page-level transcriptions.
Format:… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/CHURRO-Ottoman-Turkish-Subset.Ottoman2Turkish_dictionary[
En
The Ottoman Turkish-Turkish dictionary dataset opens the door to our cultural treasures by bringing together our centuries-old linguistic heritage with contemporary Turkish; this resource, which is accessible to everyone, removes the barriers to accessing information, democratizes learning by removing the language barrier, and allows us to freely build the bridge between the past and the present.
Tr
Osmanlıca-Türkçe sözlük veri seti, yüzlerce yıllık dil mirasımızı… See the full description on the dataset page: https://huggingface.co/datasets/TurkOpenDataOrg/Ottoman2Turkish_dictionary.ottoman_first_levelFirst level categorization of Ottoman articles.ottoman-turkish-4mUD_Ottoman_Turkish-DUDU
Summary
An Ottoman Turkish dependency treebank annotated in UD style. Created by Enes Yılandiloğlu.
Introduction
This treebank comprises 25,010 words in 2,552 sentences that are firstly automaticaly annotated via Machamp (Van der Goot et al., 2021). During the training phase, multiple modern Turkish UD treebanks were used.
and then manually corrected in a systematic way. Randomly shuffled sentences were written between 14th to 20th century in various genres
such as… See the full description on the dataset page: https://huggingface.co/datasets/enesyila/UD_Ottoman_Turkish-DUDU.Turkish-ottoman-1mottoman_second_levelSecond level categorization of Ottoman articles.ottoman-ocr-8kottoman_sentencesOttoman-DatasetOsmanlıca soru-cevap veri seti.
hed-resultsrl-dataottomanhistoryottoman
