CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01priyank-m /text_recognition_en_zh_clean Dataset Card for "text_recognition_en_zh_clean" More Information needed image1M<n<10M6 likes958 downloads4y agoHugging Face02priyank-m /MJSynth_text_recognition Dataset Card for "MJSynth_text_recognition" This is the MJSynth dataset for text recognition on document images, synthetically generated, covering 90K English words. It includes training, validation and test splits. Source of the dataset: https://www.robots.ox.ac.uk/~vgg/data/text/ Use dataset streaming functionality to try out the dataset quickly without downloading the entire dataset (refer: https://huggingface.co/docs/datasets/stream) Citation details provided on the source… See the full description on the dataset page: https://huggingface.co/datasets/priyank-m/MJSynth_text_recognition.imageimage-to-text1M<n<10M8 likes865 downloads3y agoHugging Face03priyank-m /trdg_random_en_zh_text_recognition Dataset Card for "trdg_random_en_zh_text_recognition" This synthetic dataset was generated using the TextRecognitionDataGenerator(TRDG) open source repo: https://github.com/Belval/TextRecognitionDataGenerator It contains images of text with random characters from Engilsh(en) and Chinese(zh) languages. Reference to the documentation provided by the TRDG repo: https://textrecognitiondatagenerator.readthedocs.io/en/latest/index.html imageimage-to-text100K<n<1M3 likes755 downloads2y agoHugging Face04priyank-m /text_recognition_en_zh Dataset Card for "text_recognition_en_zh" More Information needed image1M<n<10M1 likes604 downloads4y agoHugging Face05redactable-llm /synth-text-recognition Dataset Card for "Synth-Text Recognition" This is the dataset for text recognition on document images, synthetically generated, covering 90K English words. It includes training, validation and test splits. imageimage-to-text1M<n<10M4 likes575 downloads3y agoHugging Face06longhoang06 /text-recognition Dataset Card for "text-recognition" More Information needed image100K<n<1M0 likes463 downloads3y agoHugging Face07priyank-m /chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition imageimage-to-text100K<n<1M35 likes269 downloads4y agoHugging Face08Empatixx /synth-text-recognition-multilines-cs Czech Synthetic Multiline Text Recognition Dataset A large-scale synthetic dataset for Czech multiline text recognition, containing 100,000 text images with corresponding transcriptions. Created using SynthTiger. Dataset Description This dataset consists of synthetically generated images of Czech text with multiple lines per image, designed for training optical character recognition (OCR) models that can handle complex multiline text layouts. Each image contains 3 lines… See the full description on the dataset page: https://huggingface.co/datasets/Empatixx/synth-text-recognition-multilines-cs.image100K<n<1M0 likes263 downloads1y agoHugging Face09Empatixx /synth-text-recognition-cs Czech Synthetic Text Recognition Dataset A large-scale synthetic dataset for Czech text recognition, containing 454,820 text images with corresponding transcriptions. Created using SynthTiger. Dataset Description This dataset consists of synthetically generated images of Czech text, designed for training optical character recognition (OCR) models. Each image contains a single word or short phrase rendered with various visual effects to simulate real-world text appearance.… See the full description on the dataset page: https://huggingface.co/datasets/Empatixx/synth-text-recognition-cs.image100K<n<1M0 likes252 downloads1y agoHugging Face10priyank-m /IAM_words_text_recognitionimage100K<n<1M9 likes190 downloads4y agoHugging Face11priyank-m /trdg_wikipedia_en_text_recognition Dataset Card for "trdg_wikipedia_en_zh_text_recognition" This synthetic dataset was generated using the TextRecognitionDataGenerator(TRDG) open source repo: https://github.com/Belval/TextRecognitionDataGenerator It contains synthetic images of text randomly sampled from Engilsh(en) Wikipedia pages. Reference to the documentation provided by the TRDG repo: https://textrecognitiondatagenerator.readthedocs.io/en/latest/index.html imageimage-to-text100K<n<1M1 likes150 downloads2y agoHugging Face12deepcopy /text_recognition_en_zh_250k Dataset Card for "text_recognition_en_zh_250k" More Information needed image100K<n<1M0 likes116 downloads1y agoHugging Face13priyank-m /trdg_random_single_words_en_text_recognition Dataset Card for "trdg_random_single_words_en_text_recognition" More Information needed image100K<n<1M0 likes74 downloads4y agoHugging Face14sonnetechnology /license-plate-text-recognition-full Dataset Card for "license-plate-text-recognition-full" Background Information This dataset is generated from keremberke/license-plate-object-detection dataset. What we have done is: Get the Bounding Boxes for each plate in an image, Crop the image to make the plate only visible, Run it through the microsoft/trocr-large-printed model to extract the written information. Structure of the Dataset It has the same structure as the… See the full description on the dataset page: https://huggingface.co/datasets/sonnetechnology/license-plate-text-recognition-full.imageimage-to-text1K<n<10K3 likes72 downloads3y agoHugging Face15deepcopy /text_recognition_en_zh_small_250k Dataset Card for "text_recognition_en_zh_small_250k" More Information needed image100K<n<1M0 likes67 downloads1y agoHugging Face16amjad-awad /Arabic-Handwritten-Text-Recognition-Dataset Dataset Description This dataset is a re-uploaded version of the Muharaf dataset. The original dataset was created by Mehreen Saeed et al. and released for research purposes. This repository is intended for easier access and experimentation via Hugging Face. | How to use from datasets import load_dataset ds = load_dataset("amjad-awad/Arabic-Handwritten-Text-Recognition-Dataset") print(ds["train"][0]["image"]) Attribution All credit for creating and… See the full description on the dataset page: https://huggingface.co/datasets/amjad-awad/Arabic-Handwritten-Text-Recognition-Dataset.imagen<1K2 likes56 downloads9mo agoHugging Face17priyank-m /iam_sroie_text_recognitionimage100K<n<1M0 likes44 downloads4y agoHugging Face18priyank-m /trdg_dict_random_words_en_text_recognition Dataset Card for "trdg_random_words_en_text_recognition" More Information needed image100K<n<1M0 likes42 downloads4y agoHugging Face19deepcopy /handwritten-text-recognition-bongabdo Dataset Card for Bongabdo Dataset Summary Bongabdo is a curated dataset of full-page Bangla (Bengali) handwritten text, intended for use in offline handwriting recognition tasks using modern neural architectures. It includes high-resolution scanned images of handwritten Bangla scripts, transcriptions, and rich per-document metadata. The data has been contributed by people of diverse age groups, occupations, and genders, making it well-suited for training robust… See the full description on the dataset page: https://huggingface.co/datasets/deepcopy/handwritten-text-recognition-bongabdo.imagen<1K0 likes35 downloads1y agoHugging Face20napatswift /thvl_text_recognition Dataset Card for "thvl_text_recognition" More Information needed image100K<n<1M0 likes31 downloads4y agoHugging Face21priyank-m /balanced_SROIE_CHINESE_IAM_text_recognitionimage100K<n<1M0 likes27 downloads4y agoHugging Face22cytingting /chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition imageimage-to-text100K<n<1M0 likes26 downloads9mo agoHugging Face23priyank-m /word_based_IAM_SROIE_text_recognition Dataset Card for "word_based_IAM_SROIE_text_recognition" More Information needed image100K<n<1M2 likes25 downloads4y agoHugging Face24deepcopy /khmer-text-recognitionimage100K<n<1M0 likes23 downloads1y agoHugging Face25pnadel /amharic-text-recognitionimage10K<n<100K0 likes18 downloads2y agoHugging Face26Mihaiii /SROIE_2019_text_recognition-other-cols-5Subset (+ some renamings and data processing) of https://huggingface.co/datasets/priyank-m/SROIE_2019_text_recognition image1K<n<10K0 likes11 downloads2y agoHugging Face27phonsobon /khmer-text-recognitiongatedimage10K<n<100K0 likes5 downloads5mo agoHugging Face28Holmes377 /text_recognition_TextVQAEvaluate with VQA Accuracy: Note! A question might have various correct answers! So if more than 3 answers are the same, then scores 1; otherwise, scores with proportion. (here is a evaluation function that you can use) def vqa_accuracy(predictions, ground_truths_list): accuracies = [] for pred, ground_truths in zip(predictions, ground_truths_list): answer_count = collections.Counter(ground_truths) accuracy = min(1.0, answer_count[pred] / 3.0)… See the full description on the dataset page: https://huggingface.co/datasets/Holmes377/text_recognition_TextVQA.image1K<n<10K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.