datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CaptchaOCR-500K
CaptchaOCR-500K
Dataset Summary
CaptchaOCR-500K is a large-scale CAPTCHA recognition dataset containing 500,000 CAPTCHA images with corresponding text labels.
The dataset is designed for training and evaluating Optical Character Recognition (OCR), CAPTCHA solving systems, image-to-text models, and computer vision models focused on text recognition.
Tasks
Optical Character Recognition (OCR)
CAPTCHA Recognition
Image-to-Text
Computer Vision
Text… See the full description on the dataset page: https://huggingface.co/datasets/AvinashRicky/CaptchaOCR-500K.CaptchaOCR-500K
CaptchaOCR-500K
Dataset Summary
CaptchaOCR-500K is a large-scale CAPTCHA recognition dataset containing 500,000 CAPTCHA images with corresponding text labels.
The dataset is designed for training and evaluating Optical Character Recognition (OCR), CAPTCHA solving systems, image-to-text models, and computer vision models focused on text recognition.
Tasks
Optical Character Recognition (OCR)
CAPTCHA Recognition
Image-to-Text
Computer Vision
Text… See the full description on the dataset page: https://huggingface.co/datasets/dan6864/CaptchaOCR-500K.captcha-25k
Synthetic-Captcha-25k
A synthetic dataset consisting of 25,000 generated captcha images, designed for training and testing OCR and computer vision models.
Dataset Curation
Source: Generated via custom Python script (Synthetic Data).
Variety: Includes 12+ types of noise filters, distortions, and variable font rendering to simulate real-world captcha challenges.
Purpose: Created for OCR benchmarking and testing automated recognition systems.
Warning: As pure synthetic data… See the full description on the dataset page: https://huggingface.co/datasets/lumasik/captcha-25k.
