CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes11k downloads5h agoHugging Face02ybashir /CS2-HUD-OCR-Crops CS2 HUD OCR Crops Per-region HUD crops sliced from three Counter-Strike 2 match recordings, labelled where possible from the demo file's parse_ticks state. Built to train a specialist CRNN that replaces the EasyOCR killfeed reader (currently ~6 s p95 on CPU) with a sub-30 ms specialist. Source Three matches by the same POV player (farouqqq), recorded in CS2's built-in DVR + the corresponding .dem files: sample map dem rounds resolution fps sample1 Ancient… See the full description on the dataset page: https://huggingface.co/datasets/ybashir/CS2-HUD-OCR-Crops.imageimage-to-text10K<n<100K0 likes282 downloads4mo agoHugging Face03minthanthtoo-cs /Burmese-Classics-OCR-RAW Burmese (Myanmar) Books Dataset – Burmese Classics OCR (Daily Rolling Project) Overview A daily rolling dataset of Burmese books, built for AI, OCR, and NLP research.Each entry includes OCR text with metadata: title, author, page index, and Burmese character ratio. This project fills a critical gap in Burmese-language resources: Scarcity of public-domain Burmese text. High technical and financial barriers to corpus building. Enables incremental, open access for… See the full description on the dataset page: https://huggingface.co/datasets/minthanthtoo-cs/Burmese-Classics-OCR-RAW.tabular1M<n<10M1 likes189 downloads1y agoHugging Face04Reza2kn /persian-ocr-community-dataset-layout Persian OCR Community Layout Annotations Resumable layout annotations for the page images in Reza2kn/persian-ocr-community-dataset. Each row points to an exact source dataset revision, Parquet shard, blob, and row. It includes the page identifier, page dimensions, handwriting flag, and structured layout boxes produced by datalab-to/surya_layout2 at confidence threshold 0.4. The boxes field contains label, confidence, raster-order position, and pixel coordinates x0, y0, x1, y1.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-community-dataset-layout.tabularobject-detection10K<n<100K0 likes175 downloads2mo agoHugging Face05ambrosfitz /ca_1860s_ocr_v1tabular1K<n<10K0 likes168 downloads3mo agoHugging Face06tadad /kat57-ocr-bench-500-results Kat57 OCR benchmark — CER/WER Strict reference-based evaluation of 16 OCR models on a deterministic 500-card sample from Lund University Library's Kat57 catalogue-card collection. This result set contains only Character Error Rate (CER) and Word Error Rate (WER); it does not contain VLM judging or ELO ratings. The sample was drawn with seed 57 from tadad/kat57-ground-truth and is published as tadad/kat57-ground-truth-500. The OCR outputs are retained in… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-500-results.tabular1K<n<10K0 likes164 downloads21d agoHugging Face07ambrosfitz /ca_1880s_ocr_v2tabular1K<n<10K0 likes157 downloads3mo agoHugging Face08TheFinAI /MultiFinBen_OCR_Taskdocumentimage-to-text10K<n<100K0 likes149 downloads7mo agoHugging Face09tetrak /armenian-ocr-crops Tetrak Armenian OCR crops Training data for tetrak_hy, the Armenian text recogniser we are building as an EasyOCR custom model in tetrak-hy-trainer for Tetrak, an OCR pipeline for community archives. The dataset has three configurations: corpus — 1,190 proofread pages of the Armenian Soviet Encyclopedia, as plain text with full Wikisource provenance. crops — the v0 synthetic pre-training set: 181,800 rendered word crops with transcriptions. crops-v1 — the v1 synthetic training… See the full description on the dataset page: https://huggingface.co/datasets/tetrak/armenian-ocr-crops.imageimage-to-text100K<n<1M1 likes120 downloads25d agoHugging Face10ambrosfitz /ca_1830s_ocr_v2tabular1K<n<10K0 likes104 downloads3mo agoHugging Face11HeshamHaroon /arabic-turath-ocrtabular100K<n<1M0 likes100 downloads5mo agoHugging Face12ambrosfitz /ca_1860s_ocr_v2tabular1K<n<10K0 likes91 downloads3mo agoHugging Face13fatihburakkaragoz /old-nogay-turkish-ocr-corpus Old Nogay Turkish OCR Corpus This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis. We are proud to release this as a contribution to Turkish NLP, Turkic-language NLP, historical NLP, and Turkish studies. The goal is not to pretend that a small OCR corpus is a polished benchmark. The goal is more… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/old-nogay-turkish-ocr-corpus.tabulartext-generationn<1K0 likes79 downloads5mo agoHugging Face14tadad /kat57-ocr-bench-results Kat57 OCR smoke benchmark results Exact ground-truth scoring for a 50-card Tesseract integration run over tadad/kat57-ground-truth-smoke. OCR outputs are published in the tesseract config of tadad/kat57-ocr-bench. Model CER WER Evaluated Empty outputs Error sentinels Skipped references Tesseract 5 0.4656 0.8605 50 1 0 0 The corpus totals are 4,629 character edits over 9,941 reference characters and 1,221 word edits over 1,419 reference words. Scoring used… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-results.tabularn<1K0 likes70 downloads22d agoHugging Face15LLMDH /English-PD-bad-OCRtabular10K<n<100K0 likes60 downloads1y agoHugging Face16ambrosfitz /ca_1870s_ocr_v2tabular1K<n<10K0 likes58 downloads3mo agoHugging Face17alphabot2 /14_Merged_RGB_ocrThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aibot2", "total_episodes": 126, "total_frames": 54648, "total_tasks": 3, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 10, "splits": { "train": "0:126" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/14_Merged_RGB_ocr.tabularrobotics10K<n<100K0 likes53 downloads1mo agoHugging Face18ambrosfitz /ca_1870s_ocr_v1tabular1K<n<10K0 likes51 downloads3mo agoHugging Face19Lukaszl /pl-mixed-docs-ocr-dataset-100-v1-results OCR Bench Results: Polish mixed documents benchmark VLM-as-judge pairwise evaluation of OCR models on a small heterogeneous sample of Polish document-style images. Rankings depend strongly on document type, so this should be read as a document-specific OCR benchmark rather than a universal OCR ranking. This benchmark uses a lightweight 100-image Polish OCR sample covering mixed document categories such as official forms, templates, certificates, structured layouts, invoices, and… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-mixed-docs-ocr-dataset-100-v1-results.tabular1K<n<10K1 likes47 downloads6mo agoHugging Face20lkevincc0 /Akkadian-OCR-Corpus Akkadian OCR Corpus OCR-extracted corpus from Old Assyrian cuneiform tablet publications and Akkadian linguistic resources, designed for LLM pretraining on ancient Mesopotamian languages. Dataset Description This dataset contains OCR-processed text from academic publications on Old Assyrian studies, including cuneiform tablet transliterations, translations, and the Chicago Assyrian Dictionary (CAD). Splits Split Description Pages Characters OCR Quality… See the full description on the dataset page: https://huggingface.co/datasets/lkevincc0/Akkadian-OCR-Corpus.tabulartext-generation100K<n<1M0 likes43 downloads8mo agoHugging Face21davanstrien /ocr-bench-britannica-results-qwen35 OCR Bench Results: ocr-bench-britannica VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model. Leaderboard Rank Model Params ELO 95% CI Wins Losses Ties Win% 1 zai-org/GLM-OCR 0.9B 1716 1673–1769 182 60 2 75% 2 lightonai/LightOnOCR-2-1B 1B 1697 1655–1749 158 61 1 72% 3 numind/NuExtract3 4B 1649 1604–1708 162 82 0 66% 4 FireRedTeam/FireRed-OCR 2.1B 1506 1469–1548 115 127 2 47% 5… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-britannica-results-qwen35.tabularn<1K1 likes42 downloads4mo agoHugging Face22alphabot2 /Aibot2_27Aug_STEA_Pick_OCRThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aibot2", "total_episodes": 10, "total_frames": 3891, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 10, "splits": { "train": "0:10" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/Aibot2_27Aug_STEA_Pick_OCR.tabularrobotics1K<n<10K0 likes42 downloads29d agoHugging Face23alphabot2 /OCR_BIMANUAL_22_21_mergedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aibot2", "total_episodes": 89, "total_frames": 34441, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 10, "splits": { "train": "0:89" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/OCR_BIMANUAL_22_21_merged.tabularrobotics10K<n<100K0 likes37 downloads2mo agoHugging Face24ambrosfitz /ca_1890s_ocr_v2tabular1K<n<10K0 likes36 downloads3mo agoHugging Face25davanstrien /ocr-bench-judge-eval-27b OCR Bench Results: ocr-bench-britannica VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model. Leaderboard Rank Model ELO 95% CI Wins Losses Ties Win% 1 lightonai/LightOnOCR-2-1B 1675 1571–1836 26 9 1 72% 2 FireRedTeam/FireRed-OCR 1612 1518–1767 25 13 1 64% 3 zai-org/GLM-OCR 1594 1480–1739 24 14 1 62% 4 deepseek-ai/DeepSeek-OCR 1437 1332–1546 15 23 1 38% 5 rednote-hilab/dots.ocr 1182… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-judge-eval-27b.tabularn<1K0 likes33 downloads7mo agoHugging Face26alphabot2 /14_RGB_OCRThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aibot2", "total_episodes": 62, "total_frames": 25168, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 10, "splits": { "train": "0:62" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/14_RGB_OCR.tabularrobotics10K<n<100K0 likes33 downloads1mo agoHugging Face27davanstrien /bpl-ocr-bench-results OCR Bench Results: bpl-ocr-bench VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model. Leaderboard Rank Model ELO 95% CI Wins Losses Ties Win% 1 lightonai/LightOnOCR-2-1B 1559 1497–1630 39 25 0 61% 2 zai-org/GLM-OCR 1535 1471–1591 48 35 1 57% 3 rednote-hilab/dots.ocr 1453 1385–1515 26 37 0 41% 4 deepseek-ai/DeepSeek-OCR 1452 1388–1514 33 49 1 40% Details Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/bpl-ocr-bench-results.tabularn<1K0 likes32 downloads7mo agoHugging Face28harsha-desaraju /telugu-ocr-text-uncombinedtabular100K<n<1M0 likes30 downloads10mo agoHugging Face29davanstrien /ocr-bench-britannica-results OCR Bench Results: ocr-bench-britannica VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model. Leaderboard Rank Model Params ELO 95% CI Wins Losses Ties Win% 1 rednote-hilab/dots.mocr 3B 1745 1714–1782 436 141 4 75% 2 lightonai/LightOnOCR-2-1B 1B 1741 1709–1779 426 141 4 75% 3 zai-org/GLM-OCR 0.9B 1738 1707–1773 469 157 2 75% 4 allenai/olmOCR-2-7B-1025-FP8 1719 1688–1753 454 167 4 73%… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-britannica-results.tabular1K<n<10K0 likes30 downloads3mo agoHugging Face30alphabot2 /22_07_2026_OCR_BIMANUALThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aibot2", "total_episodes": 59, "total_frames": 28575, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:59" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/22_07_2026_OCR_BIMANUAL.tabularrobotics10K<n<100K0 likes30 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.