CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01harsha-desaraju /telugu-synthetic-line-imagesimage1M<n<10M0 likes1k downloads2mo agoHugging Face02harsha-desaraju /Telugu-book-text-imagesimage100K<n<1M1 likes548 downloads8mo agoHugging Face03harsha-desaraju /Telugu-text-imageimage100K<n<1M0 likes247 downloads11mo agoHugging Face04harsha-desaraju /telugu-wikisource-text-imagesimage100K<n<1M0 likes204 downloads2mo agoHugging Face05meharuhanzz /OCR-Bench1000-Telugu OCR-Bench1000-Telugu 1000 synthetic printed-text line images with ground-truth transcriptions, sampled from a larger locally-held Telugu OCR training corpus. This is a benchmark/sample release, not the full training set. Data fields Field Description file_name relative path to the image (images/...) text ground-truth transcription category telugu_only / english_only / mixed / numeric_and_symbols length_bucket short / medium / long, by character… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Telugu.imageimage-to-text1K<n<10K0 likes165 downloads10d agoHugging Face06harsha-desaraju /telugu-pdf-line-image-textimage100K<n<1M0 likes53 downloads2mo agoHugging Face07harsha-desaraju /telugu-line-ocr-bench Telugu Wikisource OCR — human-verified line crops 1044 single-line crops from Telugu Wikisource page scans, each with a transcription checked against the image by a human. Grayscale, height 64px, width a multiple of 8 — the form the encoder consumes. Columns column meaning image the line crop text gold transcription, human-verified n_graphemes akshara count of text (regex.\X) has_english text contains a Latin-script letter. Digits/punctuation do… See the full description on the dataset page: https://huggingface.co/datasets/harsha-desaraju/telugu-line-ocr-bench.imageimage-to-text1K<n<10K0 likes33 downloads2mo agoHugging Face08Divs0910 /telugu-ocr-datasetimage100K<n<1M1 likes25 downloads8mo agoHugging Face091024m /chemistry-multimodal-exams-teluguimage1K<n<10K0 likes14 downloads2y agoHugging Face10harsha-desaraju /telugu-line-text-imageimage100K<n<1M0 likes9 downloads4mo agoHugging Face11iit-patna-cse-ai /MedQA_ODD_Telugu_testimagen<1K0 likes7 downloads5mo agoHugging Face12harsha-desaraju /telugu-book-line-images-sampleimage1K<n<10K0 likes7 downloads4mo agoHugging Face13harsha-desaraju /telugu-book-line-imagesimage1M<n<10M0 likes7 downloads4mo agoHugging Face14iit-patna-cse-ai /MedQA_tran_COT_Telugu_train2gatedimage1K<n<10K0 likes5 downloads1y agoHugging Face15iit-patna-cse-ai /MedQA_tran_COT_Telugu_test2gatedimagen<1K0 likes3 downloads9mo agoHugging Face16AlbertoChestnut /telugu-ocr Telugu OCR Dataset A corpus of aligned scanned page images and human-transcribed Telugu text, sourced from Telugu Wikisource. Built for OCR model training and evaluation. Stats Total page pairs ~25,565 Books 221 Total size ~11 GB License CC BY-SA 4.0 Dataset Structure dataset/ <book_title>/ page_0001.jpg ← scan image page_0001.txt ← transcribed Telugu text (UTF-8) page_0004.jpg page_0004.txt ...… See the full description on the dataset page: https://huggingface.co/datasets/AlbertoChestnut/telugu-ocr.imageimage-to-text100K<n<1M0 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.