CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OttomanNLP /OpenITI-MAKHZAN-Ottoman-Lines Dataset Card for OpenITI MAKHZAN Ottoman Lines Dataset Summary This dataset contains line-level image-text pairs of historical Ottoman Turkish manuscripts and printed documents. It is derived from the OpenITI MAKHZAN dataset, a large aggregation of Arabic-script ground truth and evaluation data developed by the Open Islamicate Texts Initiative (OpenITI). The dataset specifically focuses on Ottoman Turkish texts and is highly valuable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OpenITI-MAKHZAN-Ottoman-Lines.imageimage-to-text1K<n<10K1 likes94 downloads3mo agoHugging Face02OttomanNLP /Akis-Ottoman-Dataset Akis-Dataset This repository provides the test dataset used in the paper "Automatic Transcription of Ottoman Documents Using Deep Learning". It contains line segment images of Ottoman documents along with their corresponding transcriptions. Dataset Overview The dataset contains 8,037 image–transcription pairs of Ottoman handwritten document line segments. Format 1: HuggingFace Dataset (Parquet — Recommended) The dataset is natively available as a… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/Akis-Ottoman-Dataset.imageimage-to-text1K<n<10K0 likes92 downloads2mo agoHugging Face03OttomanNLP /OsmanlicaEkmekveNisastaKitabi Osmanlıca Ekmek ve Nişasta Kitabı Veri Seti Bu veri seti, 1331 (1915) yılında Matbaa-i Âmire tarafından basılan, İzmir mebusu ve Dârülmuallimât sanayi-i ziraiye muallimi İhsan Bey tarafından yazılan "Kadınlara Amelî Sanayi-i Ziraiye Dersleri: Cüz 5 - Ekmek ve Nişastacılık Sanatı" kitabının görsellerini, Osmanlıca transkripsiyonunu ve (parantez içinde) günümüz Türkçesi sadeleştirmelerini/çevirilerini içerir. Veri Seti Yapısı image: Kitap sayfasının orijinal yüksek… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OsmanlicaEkmekveNisastaKitabi.imageimage-to-textn<1K0 likes60 downloads2mo agoHugging Face04OttomanNLP /CHURRO-Ottoman-Turkish-Subset CHURRO Ottoman Turkish Subset This repository contains the Ottoman Turkish historical document subset extracted from the CHURRO-DS dataset published at EMNLP 2025. CHURRO is a 3B-parameter open-weight Large Vision-Language Model (VLM) specialized for high-accuracy, low-cost historical text recognition across diverse scripts and historical variants. 📊 Dataset Overview Total Samples: 237 historical manuscript page images with page-level transcriptions. Format:… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/CHURRO-Ottoman-Turkish-Subset.imageimage-to-textn<1K0 likes44 downloads2mo agoHugging Face05asparius /ottoman-ocr-8kimage1K<n<10K1 likes11 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.