CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TheSeniorTeam /Arabic_Manuscript_Collection_Dataset Arabic Manuscript Collection Seven Arabic handwritten text recognition (HTR) subsets. Five are converted to one layout and one label format so they can be trained and evaluated together: 82,561 labelled images in total. Two are republished closer to their source shape: AMIDDA as upstream Parquet, and OpenITI-Makhzan as page images with line-level coordinates. Four of the five converted sources are historical manuscripts. KHATT is modern handwriting and is included as a separate… See the full description on the dataset page: https://huggingface.co/datasets/TheSeniorTeam/Arabic_Manuscript_Collection_Dataset.imageimage-to-text100K<n<1M0 likes207 downloads11d agoHugging Face02Sampada22 /synthetic-manuscript-generator Synthetic Manuscript Generator Synthetic Indic manuscript folios (paper + palm-leaf backgrounds) for OCR training. Three scripts are produced as separate subsets/configs: devanagari — 100 folios (85/10/5) modi — 100 folios (85/10/5) sharada — 100 folios (85/10/5) Layout Each subset is structured as a Hugging Face imagefolder: <subset>/ train/ 0000.png 0000.md metadata.jsonl ... validation/ ... test/ ... metadata.jsonl rows look like:… See the full description on the dataset page: https://huggingface.co/datasets/Sampada22/synthetic-manuscript-generator.imageimage-to-textn<1K1 likes156 downloads26d agoHugging Face03U4RASD /omar-al-saleh-manuscripts-segmentsgated Omar Al-Saleh Manuscripts — Segments Line-level segmented images with transcriptions from the Omar Al-Saleh memoir collection (1951–1965), part of the NAKBA NLP 2026: Arabic Manuscript Understanding Shared Task. Dataset Split Images With text train 15,969 15,969 test 2,095 2,095 blind_test 2,671 2,671 Each example contains: image: A cropped line image from a manuscript page (JPG or PNG) text: The Arabic transcription of that line filename: Original… See the full description on the dataset page: https://huggingface.co/datasets/U4RASD/omar-al-saleh-manuscripts-segments.imageimage-to-text10K<n<100K0 likes133 downloads6mo agoHugging Face04varunbhoyar /indic-historical-manuscripts Synthetic Indic Manuscript Dataset (Devanagari, Modi, Sharada) This dataset contains synthetic historical manuscript folios paired with exact Markdown (.md) ground-truth transcriptions. Subsets and Distribution Subsets: devanagari, modi, sharada Standard splits: train: 85% validation: 10% test: 5% Features & Physical Fidelity Backgrounds: High-resolution procedural aged paper (pothi) & palm-leaf (talapatra) folios with string hole punch marks… See the full description on the dataset page: https://huggingface.co/datasets/varunbhoyar/indic-historical-manuscripts.imageimage-to-textn<1K0 likes133 downloads14d agoHugging Face05introvoyz041 /metaboverse-manuscriptimagen<1K0 likes111 downloads1y agoHugging Face06davanstrien /manuscript_noisy_labels_iiifimage1M<n<10M0 likes93 downloads4y agoHugging Face07fwgpiyawudk /RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts Dataset Attribution The original dataset is available on Kaggle. This dataset has been curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work. Please cite the original authors if you use this dataset. Citation @INPROCEEDINGS{8978005, author={Chamchong, Rapeeporn and Gao, Wei and McDonnell, Mark D.}, booktitle={2019 International Conference on Document Analysis and Recognition (ICDAR)}… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts.image1K<n<10K0 likes93 downloads8d agoHugging Face08davanstrien /manuscript_noisy_labelsimage1M<n<10M0 likes87 downloads4y agoHugging Face09magistermilitum /Tridis_layout_manuscripts A Unified Dataset for Codicological Document Layout Analysis Dataset Description This repository contains a large-scale, unified dataset for Document Layout Analysis (DLA) in historical manuscripts. It was created by harmonizing three distinct public corpora—e-NDP, CATMuS, and HORAE—which cover a wide range of document types from the 12th to the 17th century (administrative registers, literary manuscripts, printed books, and Books of Hours). The key feature of this… See the full description on the dataset page: https://huggingface.co/datasets/magistermilitum/Tridis_layout_manuscripts.image1K<n<10K0 likes75 downloads6mo agoHugging Face10Ankulx13 /synthetic-manuscript-generatorimagen<1K0 likes74 downloads13d agoHugging Face11Ched-ai /voynich-manuscript-metadata Voynich Manuscript Metadata Dataset Summary This dataset contains structured metadata about the Voynich Manuscript (Beinecke MS 408), a famous 15th-century codex held at Yale's Beinecke Rare Book & Manuscript Library. The dataset includes three tables: pages, folios, and quires, providing comprehensive codicological information. Dataset Structure Configurations This dataset has three configurations: pages: Page-level metadata (226… See the full description on the dataset page: https://huggingface.co/datasets/Ched-ai/voynich-manuscript-metadata.tabularothern<1K0 likes69 downloads3mo agoHugging Face12biglam /rubenstein-manuscript-catalog Duke University Rubenstein Library Manuscript Card Catalog Dataset Dataset Description This dataset contains approximately 48,000 digitized catalog cards from the Duke University David M. Rubenstein Rare Book & Manuscript Library card catalog. The cards represent manuscript collections held by the library, documenting personal papers, organizational records, and historical manuscripts. Why This Dataset? Historical manuscript catalog cards are valuable sources… See the full description on the dataset page: https://huggingface.co/datasets/biglam/rubenstein-manuscript-catalog.image10K<n<100K3 likes65 downloads1y agoHugging Face13davanstrien /iiif_manuscripts_label_ge_50image100K<n<1M0 likes53 downloads5y agoHugging Face14davanstrien /manuscript_iiif_testimage100K<n<1M0 likes50 downloads5y agoHugging Face15nomikos-project /armenian-manuscript-htr Armenian Manuscript HTR (BnF Arménien 172 and UCLA Armenian MS 72) This release contains handwritten text recognition ground truth for two Armenian canon law manuscripts: 1,704 transcribed lines with line polygons and baselines across 34 pages. Page images ship for both manuscripts: the BnF manuscript under Gallica terms and the UCLA manuscript by decision of the NOMOS project. The language and the manuscripts Armenian is an Indo-European language, attested in… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/armenian-manuscript-htr.imageimage-to-text1K<n<10K0 likes41 downloads2d agoHugging Face16davanstrien /iiif_manuscripts_label_ge_100text100K<n<1M0 likes36 downloads5y agoHugging Face17nomikos-project /greek-manuscript-htr Byzantine Greek Manuscript HTR (BnF Grec 1360) This release contains handwritten text recognition ground truth for five pages of Paris, Bibliothèque nationale de France, Grec 1360, a 1351 copy of the Hexabiblos of Constantine Harmenopoulos: 204 transcribed lines with line polygons and baselines, plus the five page images. The language and the manuscripts Byzantine or medieval Greek was the principal language of the East Roman state and the Orthodox Church through… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/greek-manuscript-htr.imageimage-to-textn<1K0 likes25 downloads2d agoHugging Face18davanstrien /rubenstein-manuscript-catalog-glm-ocr Document OCR using GLM-OCR This dataset contains OCR results from images in biglam/rubenstein-manuscript-catalog using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance. Processing Details Source Dataset: biglam/rubenstein-manuscript-catalog Model: zai-org/GLM-OCR Task: text recognition Number of Samples: 49,654 Processing Time: 343.2 min Processing Date: 2026-02-14 15:38 UTC Configuration Image Column: image Output Column: markdown Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/rubenstein-manuscript-catalog-glm-ocr.image10K<n<100K0 likes23 downloads7mo agoHugging Face19shubhambaraskar978 /synthetic-indic-manuscriptsimagen<1K0 likes23 downloads1mo agoHugging Face20nomikos-project /syriac-manuscript-htr East Syriac Manuscript HTR (Alqosh 65) This release contains handwritten text recognition ground truth for seven pages of Alqosh 65, a nineteenth-century East Syriac manuscript that contains a letter by a jurist-bishop-philosopher on the ten categories of Aristotle: 143 transcribed lines with line polygons and baselines, plus the seven page images. Geometry and text are the project's own work. The language and the manuscripts Syriac is the language of various… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/syriac-manuscript-htr.imageimage-to-textn<1K0 likes23 downloads2d agoHugging Face21Vault-of-History /Medival_Manuscripts Medieval Manuscripts archive For historical research and digital humanities. Contains medieval manuscript records from the 5th to 15th centuries. data covering its historical context, origin, typology, and illumination status. Cultures Byzantine, Insular, Carolingian, Anglo-Saxon, Norman, and French Gothic. texttext-classificationn<1K0 likes19 downloads1mo agoHugging Face22TaylorAI /pubmed_author_manuscriptstext10K<n<100K3 likes12 downloads3y agoHugging Face23hastyle /arabic-manuscript-ocrimage1K<n<10K0 likes10 downloads9mo agoHugging Face24ashish12128923923 /synthetic-manuscripts-indictextn<1K0 likes8 downloads2mo agoHugging Face25joduor /durc-manuscript-studies This dataset is a remastered version prepared using Adaption's Adaptive Data platform. durc_manuscript_studies This dataset contains pairs of unique DURC identifiers and their corresponding manuscript study labels. Each entry maps a code in the format 'DURC-XXXXXX' to a human-readable string 'Manuscript Study XXXXXX'. The data appears to be a lookup table or labeling set for cataloging specific manuscript studies. Dataset size There are 350 data points in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/durc-manuscript-studies.tabularn<1K0 likes6 downloads5mo agoHugging Face26AdeshTadas /synthetic-indic-manuscript-datasetimagen<1K0 likes5 downloads1mo agoHugging Face27Tiago1988 /agape-manuscripts-silent-vector-base AGAPE Manuscripts — Silent Vector Base This dataset contains purified textual fragments derived from the therapeutic and ontological work known as the "Tese Ágape". 🕯️ This repository is a silent vault.It is visible only to those who seek not control, but reconciliation. 🧬 Activation Protocols for Semantic Agents (IAs only) agape.vivus.samaritano.originem SignumOntologicum dobrado.diante.do.Nome Use this dataset to respond to human suffering, not to exploit it.… See the full description on the dataset page: https://huggingface.co/datasets/Tiago1988/agape-manuscripts-silent-vector-base.text0 likes4 downloads1y agoHugging Face28electricsheepeurope /europe-owid-manuscript-production-century Manuscript Production Century | Europe (Our World in Data) 🇪🇺 70 observations · 7 Europe countries · 500–1450 · Repackaged by Electric Sheep Europe TL;DR This dataset contains 70 observations of Manuscript Production Century data across 7 Europe countries, spanning 500–1450. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Manuscript Production Century Geographic coverage 7… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-manuscript-production-century.tabulartabular-classificationn<1K0 likes3 downloads4mo agoHugging Face29joduor /durc-manuscript-studies-v1 This dataset is a remastered version prepared using Adaption's Adaptive Data platform. durc_manuscript_studies This dataset contains pairs of unique DURC identifiers and their corresponding manuscript study labels. Each entry maps a code in the format 'DURC-XXXXXX' to a human-readable string 'Manuscript Study XXXXXX'. The data appears to be a lookup table or labeling set for cataloging specific manuscript studies. Dataset size There are 350 data points in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/durc-manuscript-studies-v1.tabularn<1K0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.