datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenITI-MAKHZAN-Ottoman-Lines
Dataset Card for OpenITI MAKHZAN Ottoman Lines
Dataset Summary
This dataset contains line-level image-text pairs of historical Ottoman Turkish manuscripts and printed documents. It is derived from the OpenITI MAKHZAN dataset, a large aggregation of Arabic-script ground truth and evaluation data developed by the Open Islamicate Texts Initiative (OpenITI).
The dataset specifically focuses on Ottoman Turkish texts and is highly valuable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OpenITI-MAKHZAN-Ottoman-Lines.Akis-Ottoman-Dataset
Akis-Dataset
This repository provides the test dataset used in the paper "Automatic Transcription of Ottoman Documents Using Deep Learning". It contains line segment images of Ottoman documents along with their corresponding transcriptions.
Dataset Overview
The dataset contains 8,037 image–transcription pairs of Ottoman handwritten document line segments.
Format 1: HuggingFace Dataset (Parquet — Recommended)
The dataset is natively available as a… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/Akis-Ottoman-Dataset.OsmanlicaEkmekveNisastaKitabi
Osmanlıca Ekmek ve Nişasta Kitabı Veri Seti
Bu veri seti, 1331 (1915) yılında Matbaa-i Âmire tarafından basılan, İzmir mebusu ve Dârülmuallimât sanayi-i ziraiye muallimi İhsan Bey tarafından yazılan "Kadınlara Amelî Sanayi-i Ziraiye Dersleri: Cüz 5 - Ekmek ve Nişastacılık Sanatı" kitabının görsellerini, Osmanlıca transkripsiyonunu ve (parantez içinde) günümüz Türkçesi sadeleştirmelerini/çevirilerini içerir.
Veri Seti Yapısı
image: Kitap sayfasının orijinal yüksek… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OsmanlicaEkmekveNisastaKitabi.CHURRO-Ottoman-Turkish-Subset
CHURRO Ottoman Turkish Subset
This repository contains the Ottoman Turkish historical document subset extracted from the CHURRO-DS dataset published at EMNLP 2025.
CHURRO is a 3B-parameter open-weight Large Vision-Language Model (VLM) specialized for high-accuracy, low-cost historical text recognition across diverse scripts and historical variants.
📊 Dataset Overview
Total Samples: 237 historical manuscript page images with page-level transcriptions.
Format:… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/CHURRO-Ottoman-Turkish-Subset.ottoman-ocr-8k
