CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BDRC /monlamai-transcriptions Tibetan OCR — MonlamAI transcriptions 3,072 page images of Tibetan dbu-med (u-med) manuscripts with page-level Unicode transcriptions, contributed by MonlamAI over BDRC manuscript scans and aligned page by page. This is the full page equivalent of the line-segmented datasets available on openpecha/OCR-Betsug and openpecha/OCR-Drutsa, with the following changes: filter out cases where not all line in a page are transcribed (so not all the lines in the original datasets are… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/monlamai-transcriptions.imageimage-to-text1K<n<10K0 likes59 downloads1mo agoHugging Face02BDRC /berkeley-transcriptions Tibetan OCR — Berkeley 8,866 page images of Tibetan text with page-level Unicode transcriptions, mostly dbu-med (u-med) manuscripts with a small woodblock (uchen) portion. These works were transcribed by Geshe Dangsong Namgyal, from 2016 to 2026, in his role as a Data System Analyst at University of California, Berkeley. This work was initiated by Prof. Kurt Keutzer with the longstanding hope of producing training data for Handwritten Text Recognition for dbu med and OCR of… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/berkeley-transcriptions.imageimage-to-text1K<n<10K0 likes57 downloads1mo agoHugging Face03BDRC /palri-parkhang-transcriptions Tibetan OCR — Palri Parkhang 11,133 page images of Tibetan text with page-level Unicode transcriptions, mostly dbu-med (u-med) manuscripts with a small uchen portion. The transcriptions were produced by Palri Parkhang, an input project led by Chris Tomlinson (former BDRC's CTO) in Nepal in 2006-2013 and aligned to BDRC scans; the material is largely Nyingma collected works and gter ma cycles. The transcriptions prioritized legibility over fidelity to the enscribed text and… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/palri-parkhang-transcriptions.imageimage-to-text10K<n<100K0 likes56 downloads1mo agoHugging Face04BSC-CSSH /AMSMB-line-transcription Dataset Card Dataset for line-level handwritten text recognition on medieval historical manuscripts, consisting of 3,369 lines (images of text lines with the associated transcription and metadata) from 100 digitized documents written by at least 80 different hands and spanning three centuries (from 1208 to 1499). This dataset is derived from the AMSMB dataset, which contains the full-page images of the digitized manuscripts and their associated transcriptions in the PageXML format.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-CSSH/AMSMB-line-transcription.imageimage-to-text1K<n<10K0 likes32 downloads1y agoHugging Face05justinsunqiu /multilingual_transcriptions_summarized_by_english_backtranslated_finalimage10K<n<100K0 likes21 downloads1y agoHugging Face06justinsunqiu /multilingual_transcriptions_finalimage1K<n<10K0 likes13 downloads1y agoHugging Face07justinsunqiu /multilingual_transcriptions_fullimage1K<n<10K0 likes10 downloads1y agoHugging Face08justinsunqiu /multilingual_transcriptions_translated_rawimage1K<n<10K0 likes9 downloads1y agoHugging Face09justinsunqiu /multilingual_transcriptions_cleanedimage1K<n<10K0 likes8 downloads1y agoHugging Face10justinsunqiu /multilingual_transcriptions_summarized_by_native_nonnativeimage1K<n<10K0 likes8 downloads1y agoHugging Face11justinsunqiu /multilingual_transcriptions_translated_english_finalimage1K<n<10K0 likes8 downloads1y agoHugging Face12justinsunqiu /multilingual_transcriptions_summarizedimage1K<n<10K0 likes7 downloads1y agoHugging Face13justinsunqiu /multilingual_transcriptions_summarized_by_type_finalimage1K<n<10K0 likes6 downloads1y agoHugging Face14justinsunqiu /multilingual_transcriptionsimage1K<n<10K0 likes5 downloads2y agoHugging Face15justinsunqiu /transcription_changesimagen<1K0 likes5 downloads1y agoHugging Face16justinsunqiu /transcription_changes_classifiedimagen<1K0 likes5 downloads1y agoHugging Face17justinsunqiu /multilingual_transcriptions_rawimage1K<n<10K0 likes3 downloads2y agoHugging Face18curiousmrk /transcription-coding-wiki-500kgated Transcription Dataset: Code & Wiki (390K) Text-to-image rendered dataset for training vision-language models to read code and text from images. Schema Column Type Description image Image Rendered grayscale JPEG prompt string Transcription instruction (varied) response string Ground truth text language string python/javascript/java/c++/rust/go/english domain string code or english length_bucket string short/medium/long/gundam resolution string… See the full description on the dataset page: https://huggingface.co/datasets/curiousmrk/transcription-coding-wiki-500k.image100K<n<1M0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.