datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CaptchaOCR-500K
CaptchaOCR-500K
Dataset Summary
CaptchaOCR-500K is a large-scale CAPTCHA recognition dataset containing 500,000 CAPTCHA images with corresponding text labels.
The dataset is designed for training and evaluating Optical Character Recognition (OCR), CAPTCHA solving systems, image-to-text models, and computer vision models focused on text recognition.
Tasks
Optical Character Recognition (OCR)
CAPTCHA Recognition
Image-to-Text
Computer Vision
Text… See the full description on the dataset page: https://huggingface.co/datasets/AvinashRicky/CaptchaOCR-500K.Deep_Captcha
DeepCaptcha Dataset: AI-Resistant CAPTCHA Benchmark
Overview
DeepCaptcha is a professional-grade dataset designed for benchmarking the AI-resistance and robustness of computer vision models. It contains CAPTCHA images generated with varying levels of adversarial protection, specifically engineered to thwart common automated recognition attacks (CNNs, OCR, etc.) while remaining human-readable.
Dataset Structure
The dataset is organized into 5 primary difficulty… See the full description on the dataset page: https://huggingface.co/datasets/Knight07/Deep_Captcha.Open-Captcha-Image-DLC
File Breakdown:
File Name
Size
Description
.gitattributes
2.46 kB
Rules for Git Large File Storage (LFS).
README.md
42 Bytes
Basic project description.
Usage:
This dataset can be used for:
Training CAPTCHA Solvers:Build models capable of solving image-based CAPTCHA challenges.
Testing Automation:Evaluate CAPTCHA-breaking algorithms for robustness.
Security Research:Understand the limitations of CAPTCHA systems and enhance security protocols.
captcha-25k
Synthetic-Captcha-25k
A synthetic dataset consisting of 25,000 generated captcha images, designed for training and testing OCR and computer vision models.
Dataset Curation
Source: Generated via custom Python script (Synthetic Data).
Variety: Includes 12+ types of noise filters, distortions, and variable font rendering to simulate real-world captcha challenges.
Purpose: Created for OCR benchmarking and testing automated recognition systems.
Warning: As pure synthetic data… See the full description on the dataset page: https://huggingface.co/datasets/lumasik/captcha-25k.CaptchaOCR-500K
CaptchaOCR-500K
Dataset Summary
CaptchaOCR-500K is a large-scale CAPTCHA recognition dataset containing 500,000 CAPTCHA images with corresponding text labels.
The dataset is designed for training and evaluating Optical Character Recognition (OCR), CAPTCHA solving systems, image-to-text models, and computer vision models focused on text recognition.
Tasks
Optical Character Recognition (OCR)
CAPTCHA Recognition
Image-to-Text
Computer Vision
Text… See the full description on the dataset page: https://huggingface.co/datasets/dan6864/CaptchaOCR-500K.sat-captchas-v1
SAT Captchas V1
This dataset archive contains CAPTCHA images from the SAT CFDI verification flow targeted by the sat-captcha-solver project.
Release Context
As of the morning of March 25, 2026, the specific CAPTCHA targeted by this project had been deprecated. That deprecation is the reason the related code and this dataset archive were made public.
Contents
images.tar: archive containing the CAPTCHA image files under images/
labels.csv: filename-to-label… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/sat-captchas-v1.
