CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ved1245 /synthetic-manuscript-dataset Synthetic Manuscript Dataset Synthetic historical manuscript folios generated using an automated Python pipeline. Scripts The dataset contains three script configurations: Devanagari Modi Sharada Each script contains 100 synthetic manuscript folios. Dataset Splits Split Samples per Script Train 85 Validation 10 Test 5 Total 100 Across all three scripts, the dataset contains: 300 manuscript images 300 corresponding Markdown… See the full description on the dataset page: https://huggingface.co/datasets/ved1245/synthetic-manuscript-dataset.imagen<1K1 likes628 downloads14d agoHugging Face02PrachitiKothekar /synthetic-manuscriptsimagen<1K0 likes353 downloads29d agoHugging Face03Shubhamyadav321 /synthetic-manuscript-generatorimagen<1K0 likes236 downloads8d agoHugging Face04QFun /MANUS-HaGRID MANUS-HaGRID: HaGRID-derived Multimodal Annotated Naturalistic Hand Understanding Dataset MANUS-HaGRID is the HaGRID/HaGRIDv2-derived subset of the Multimodal Annotated Naturalistic Hand Understanding (MANUS) dataset family. It provides multimodal annotations for naturalistic hand gesture understanding, including RGB images, hand crops, depth maps, 2D bounding boxes, estimated MANO-style hand mesh metadata, and multi-view mesh renderings where available. This repository contains… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS-HaGRID.imageimage-to-image10K<n<100K0 likes227 downloads3mo agoHugging Face05TheSeniorTeam /Arabic_Manuscript_Collection_Dataset Arabic Manuscript Collection Seven Arabic handwritten text recognition (HTR) subsets. Five are converted to one layout and one label format so they can be trained and evaluated together: 82,561 labelled images in total. Two are republished closer to their source shape: AMIDDA as upstream Parquet, and OpenITI-Makhzan as page images with line-level coordinates. Four of the five converted sources are historical manuscripts. KHATT is modern handwriting and is included as a separate… See the full description on the dataset page: https://huggingface.co/datasets/TheSeniorTeam/Arabic_Manuscript_Collection_Dataset.imageimage-to-text100K<n<1M0 likes209 downloads13d agoHugging Face06biglam /yalta_ai_segmonto_manuscript_dataset YALTAi SegmOnto Manuscript and Early Printed Book Dataset 1,147 page images from manuscripts and early printed books, 9th to 17th century, with bounding-box annotations for zone types drawn from the SegmOnto vocabulary. Created by Thibault Clérice and deposited on Zenodo alongside the paper You Actually Look Twice At it (YALTAi) (Journal of Data Mining and Digital Humanities, 2022), which treats page layout recognition on historical documents as an object detection problem… See the full description on the dataset page: https://huggingface.co/datasets/biglam/yalta_ai_segmonto_manuscript_dataset.imageobject-detection1K<n<10K2 likes175 downloads2mo agoHugging Face07Sampada22 /synthetic-manuscript-generator Synthetic Manuscript Generator Synthetic Indic manuscript folios (paper + palm-leaf backgrounds) for OCR training. Three scripts are produced as separate subsets/configs: devanagari — 100 folios (85/10/5) modi — 100 folios (85/10/5) sharada — 100 folios (85/10/5) Layout Each subset is structured as a Hugging Face imagefolder: <subset>/ train/ 0000.png 0000.md metadata.jsonl ... validation/ ... test/ ... metadata.jsonl rows look like:… See the full description on the dataset page: https://huggingface.co/datasets/Sampada22/synthetic-manuscript-generator.imageimage-to-textn<1K1 likes156 downloads29d agoHugging Face08U4RASD /omar-al-saleh-manuscripts-segmentsgated Omar Al-Saleh Manuscripts — Segments Line-level segmented images with transcriptions from the Omar Al-Saleh memoir collection (1951–1965), part of the NAKBA NLP 2026: Arabic Manuscript Understanding Shared Task. Dataset Split Images With text train 15,969 15,969 test 2,095 2,095 blind_test 2,671 2,671 Each example contains: image: A cropped line image from a manuscript page (JPG or PNG) text: The Arabic transcription of that line filename: Original… See the full description on the dataset page: https://huggingface.co/datasets/U4RASD/omar-al-saleh-manuscripts-segments.imageimage-to-text10K<n<100K0 likes139 downloads6mo agoHugging Face09varunbhoyar /indic-historical-manuscripts Synthetic Indic Manuscript Dataset (Devanagari, Modi, Sharada) This dataset contains synthetic historical manuscript folios paired with exact Markdown (.md) ground-truth transcriptions. Subsets and Distribution Subsets: devanagari, modi, sharada Standard splits: train: 85% validation: 10% test: 5% Features & Physical Fidelity Backgrounds: High-resolution procedural aged paper (pothi) & palm-leaf (talapatra) folios with string hole punch marks… See the full description on the dataset page: https://huggingface.co/datasets/varunbhoyar/indic-historical-manuscripts.imageimage-to-textn<1K0 likes133 downloads16d agoHugging Face10davanstrien /manuscript_noisy_labelsimage1M<n<10M0 likes110 downloads4y agoHugging Face11davanstrien /manuscript_noisy_labels_iiifimage1M<n<10M0 likes107 downloads4y agoHugging Face12introvoyz041 /metaboverse-manuscriptimagen<1K0 likes98 downloads1y agoHugging Face13fwgpiyawudk /RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts Dataset Attribution The original dataset is available on Kaggle. This dataset has been curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work. Please cite the original authors if you use this dataset. Citation @INPROCEEDINGS{8978005, author={Chamchong, Rapeeporn and Gao, Wei and McDonnell, Mark D.}, booktitle={2019 International Conference on Document Analysis and Recognition (ICDAR)}… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts.image1K<n<10K0 likes96 downloads11d agoHugging Face14kamkarprince /synthetic-indic-manuscriptsimageimage-to-textn<1K1 likes95 downloads5d agoHugging Face15QFun /MANUS-DexYCB MANUS-DexYCB: DexYCB-derived Multimodal Annotated Naturalistic Hand Understanding Dataset MANUS-DexYCB is the DexYCB-derived subset of the Multimodal Annotated Naturalistic Hand Understanding (MANUS) dataset family. It is released as a source-specific repository because MANUS subsets are governed by different upstream licenses. This repository contains only the DexYCB-derived MANUS test split. HaGRID/HaGRIDv2-derived data is released separately as MANUS-HaGRID.… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS-DexYCB.textimage-to-image0 likes80 downloads5mo agoHugging Face16Ankulx13 /synthetic-manuscript-generatorimagen<1K0 likes75 downloads16d agoHugging Face17biglam /rubenstein-manuscript-catalog Duke University Rubenstein Library Manuscript Card Catalog Dataset Dataset Description This dataset contains approximately 48,000 digitized catalog cards from the Duke University David M. Rubenstein Rare Book & Manuscript Library card catalog. The cards represent manuscript collections held by the library, documenting personal papers, organizational records, and historical manuscripts. Why This Dataset? Historical manuscript catalog cards are valuable sources… See the full description on the dataset page: https://huggingface.co/datasets/biglam/rubenstein-manuscript-catalog.image10K<n<100K3 likes62 downloads1y agoHugging Face18magistermilitum /Tridis_layout_manuscripts A Unified Dataset for Codicological Document Layout Analysis Dataset Description This repository contains a large-scale, unified dataset for Document Layout Analysis (DLA) in historical manuscripts. It was created by harmonizing three distinct public corpora—e-NDP, CATMuS, and HORAE—which cover a wide range of document types from the 12th to the 17th century (administrative registers, literary manuscripts, printed books, and Books of Hours). The key feature of this… See the full description on the dataset page: https://huggingface.co/datasets/magistermilitum/Tridis_layout_manuscripts.image1K<n<10K0 likes59 downloads6mo agoHugging Face19davanstrien /manuscript_iiif_testimage100K<n<1M0 likes50 downloads5y agoHugging Face20nomikos-project /armenian-manuscript-htr Armenian Manuscript HTR (BnF Arménien 172 and UCLA Armenian MS 72) This release contains handwritten text recognition ground truth for two Armenian canon law manuscripts: 1,704 transcribed lines with line polygons and baselines across 34 pages. Page images ship for both manuscripts: the BnF manuscript under Gallica terms and the UCLA manuscript by decision of the NOMOS project. The language and the manuscripts Armenian is an Indo-European language, attested in… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/armenian-manuscript-htr.imageimage-to-text1K<n<10K0 likes50 downloads5d agoHugging Face21davanstrien /iiif_manuscripts_label_ge_50image100K<n<1M0 likes45 downloads5y agoHugging Face22prachibangre100 /synthetic-manuscript-datasetimagen<1K0 likes41 downloads1mo agoHugging Face23mmargenot /bph_manuscript_collectionThe Bibliotheca Philosophica Hermetica is a library containing a broad collection of books on Hermetic philosophy, magic, and the occult. This dataset is a subset of the books from that library: Filtered down to books in english Transcribed into text files Stored in JSON with a mapping of filename -> content for each page. These books are from the scanned collection at the Embassy of the Free Mind and are from the early 20th century and before. imagetranslation1K<n<10K0 likes40 downloads1y agoHugging Face24nomikos-project /greek-manuscript-htr Byzantine Greek Manuscript HTR (BnF Grec 1360) This release contains handwritten text recognition ground truth for five pages of Paris, Bibliothèque nationale de France, Grec 1360, a 1351 copy of the Hexabiblos of Constantine Harmenopoulos: 204 transcribed lines with line polygons and baselines, plus the five page images. The language and the manuscripts Byzantine or medieval Greek was the principal language of the East Roman state and the Orthodox Church through… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/greek-manuscript-htr.imageimage-to-textn<1K0 likes30 downloads5d agoHugging Face25nomikos-project /syriac-manuscript-htr East Syriac Manuscript HTR (Alqosh 65) This release contains handwritten text recognition ground truth for seven pages of Alqosh 65, a nineteenth-century East Syriac manuscript that contains a letter by a jurist-bishop-philosopher on the ten categories of Aristotle: 143 transcribed lines with line polygons and baselines, plus the seven page images. Geometry and text are the project's own work. The language and the manuscripts Syriac is the language of various… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/syriac-manuscript-htr.imageimage-to-textn<1K0 likes28 downloads5d agoHugging Face26QFun /MANUS MANUS Dataset Family Multimodal Annotated Naturalistic Hand Understanding (MANUS) is a family of source-specific multimodal hand-understanding dataset releases for naturalistic hand analysis, hand-aware image generation, 2D/3D hand understanding, depth estimation, and multi-view hand representation learning. MANUS does not use a single dataset-wide license. Each source-specific release is governed by its own upstream-compatible license. This landing repository indexes the available… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS.textn<1K0 likes24 downloads5mo agoHugging Face27davanstrien /rubenstein-manuscript-catalog-glm-ocr Document OCR using GLM-OCR This dataset contains OCR results from images in biglam/rubenstein-manuscript-catalog using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance. Processing Details Source Dataset: biglam/rubenstein-manuscript-catalog Model: zai-org/GLM-OCR Task: text recognition Number of Samples: 49,654 Processing Time: 343.2 min Processing Date: 2026-02-14 15:38 UTC Configuration Image Column: image Output Column: markdown Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/rubenstein-manuscript-catalog-glm-ocr.image10K<n<100K0 likes23 downloads7mo agoHugging Face28ahmadAlrabghi /manusctips_data_for_OCR_splittedimagen<1K0 likes21 downloads1y agoHugging Face29Manusinhh /cxr-10k-reasoning-dataset 🫁 CXR-10K Reasoning Dataset A dataset of 10,000 chest X-ray images paired with step-by-step clinical reasoning and radiology impression summaries, curated for training and evaluating medical vision-language models like MedGEMMA, LLaVA-Med, and others. 📂 Dataset Structure This dataset is saved in Arrow format and was built using the Hugging Face datasets library. Each sample includes: image: Chest X-ray image (PNG or JPEG) reasoning: Step-wise radiological reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/Manusinhh/cxr-10k-reasoning-dataset.image10K<n<100K1 likes20 downloads1y agoHugging Face30shubhambaraskar978 /synthetic-indic-manuscriptsimagen<1K0 likes15 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.