datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VenusBench-CAPTCHA
VenusBench-CAPTCHA: A Real-World CAPTCHA Screenshot–Action Benchmark for GUI Agents
Evaluation Code: https://github.com/inclusionAI/UI-Venus/tree/VenusBench-CAPTCHA
Introduction
CAPTCHA solving is a practical challenge for multimodal GUI agents because it requires more than isolated visual recognition. An agent must understand the challenge instruction, identify the relevant interface region, recognize or reason about the visual target, ground the result… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/VenusBench-CAPTCHA.captcha-like-suicaOpen_CaptchaWorld
Open CaptchaWorld Dataset
This dataset accompanies the paper Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents. It contains 20 distinct CAPTCHA types, each testing different visual reasoning capabilities. The dataset is designed for evaluating the visual reasoning and interaction capabilities of Multimodal Large Language Model (MLLM)-powered agents.
Project Page | Github
The dataset includes:
20 CAPTCHA Types: A diverse set of… See the full description on the dataset page: https://huggingface.co/datasets/OpenCaptchaWorld/Open_CaptchaWorld.synthetic-captchas-library
🌍 Synthetic Multilingual CAPTCHA Library
Repository: remiai3/synthetic-captchas-libraryA multilingual dataset of synthetic 4-character CAPTCHA images designed for OCR, multilingual vision models, and script recognition research.
This dataset spans 44 world writing systems and is especially useful for low-resource script OCR training.
📌 Dataset Summary
Each script includes 100,000 unique CAPTCHA images.The dataset is provided in two parallel formats:
CSV version… See the full description on the dataset page: https://huggingface.co/datasets/remiai3/synthetic-captchas-library.captcha-dataCaptchaCaptchascaptcha-imagesCaptcha images dataset.captcha-datatext-captcha-data-cleanCaptchaOCR-500K
CaptchaOCR-500K
Dataset Summary
CaptchaOCR-500K is a large-scale CAPTCHA recognition dataset containing 500,000 CAPTCHA images with corresponding text labels.
The dataset is designed for training and evaluating Optical Character Recognition (OCR), CAPTCHA solving systems, image-to-text models, and computer vision models focused on text recognition.
Tasks
Optical Character Recognition (OCR)
CAPTCHA Recognition
Image-to-Text
Computer Vision
Text… See the full description on the dataset page: https://huggingface.co/datasets/AvinashRicky/CaptchaOCR-500K.captcha
best.pth
model = CaptchaCNNTransformer(
img_h=32,
img_w=128,
dim=256,
depth=6,
heads=4,
num_classes=num_classes,
channels=1,
dropout=0.2,
).to(device)
phpwind-captcha-dataset
PHPWind Captcha Dataset
Labelled four-digit numeric captcha images for training and evaluating OCR on
authorized PHPWind deployments. This repository stores each visual captcha
family in an independent dataset directory so that samples from different
versions, forks, themes, or generators are never silently mixed.
中文说明: README_zh.md
Related model: FlanChanXwO/phpwind-captcha-ocr
Dataset catalogue
Dataset ID
Deployment or version identifier
Status
Images… See the full description on the dataset page: https://huggingface.co/datasets/FlanChanXwO/phpwind-captcha-dataset.Open_CaptchaWorld
Open CaptchaWorld Dataset
This dataset accompanies the paper Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents. It contains 20 distinct CAPTCHA types, each testing different visual reasoning capabilities. The dataset is designed for evaluating the visual reasoning and interaction capabilities of Multimodal Large Language Model (MLLM)-powered agents.
Project Page | Github
The dataset includes:
20 CAPTCHA Types: A diverse set of… See the full description on the dataset page: https://huggingface.co/datasets/Medli1108/Open_CaptchaWorld.Open_CaptchaWorld
Open CaptchaWorld Dataset
This dataset accompanies the paper Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents. It contains 20 distinct CAPTCHA types, each testing different visual reasoning capabilities. The dataset is designed for evaluating the visual reasoning and interaction capabilities of Multimodal Large Language Model (MLLM)-powered agents.
Project Page | Github
The dataset includes:
20 CAPTCHA Types: A diverse set of… See the full description on the dataset page: https://huggingface.co/datasets/YaxinLuo/Open_CaptchaWorld.NextGen-CAPTCHAs4-captcha
4-captcha dataset
Grayscale synthetic four-digit captcha images for robust digit-string recognition and adversarial fine-tuning experiments.
Splits
Split
Images
Labels
clean/train
100,000
4-digit string
clean/val
5,000
4-digit string
clean/test
5,000
4-digit string
adv/{vit,cnn}/train
20,000 each
same as source clean image
adv/{vit,cnn}/val
1,000 each
same as source clean image
adv/{vit,cnn}/test
1,000 each
same as source clean image… See the full description on the dataset page: https://huggingface.co/datasets/pymlex/4-captcha.funtime_captchaFunTime Captcha 22.01.26
captcha-dataset
Synthetic CAPTCHA Dataset
Synthetic distorted-text CAPTCHA images with exact character-string labels, used
to fine-tune Hansda/minicpm-v4_6-captcha-lora.
Contents
Format: HuggingFace imagefolder (image, text columns), with
metadata.jsonl per split.
Splits: train (10,000) and test (2,000) — 12,000 images total.
Labels: 4–7 characters from ABCDEFGHJKLMNPQRSTUVWXYZabcdefghjkmnpqrstuvwxyz23456789
(visually ambiguous I O 0 1 l are excluded).
Generation… See the full description on the dataset page: https://huggingface.co/datasets/Hansda/captcha-dataset.length_captchas
Length Captchas Dataset
Converted dataset from twitter_ai label tool.
maplestory_captchaA huge collection of English MapleStory's captcha text in jpg that I have collected over the years. ENJOY!!
It us used by pre-Big Bang MapleStory, throughout the game from Lie-Detector (anti-macro item), logins, to NPC conversations.
Up till version 190 when they have switched using Runes (Up, Down, Left, Right arrow keys) for most of the time for detection of macros and bots.
These images are not labelled, I'm releasing this for anyone that wants the dataset to be able to train a model… See the full description on the dataset page: https://huggingface.co/datasets/lastbattle/maplestory_captcha.captcha_100kcaptcha-to-text
eKYB Captcha Labeled Dataset
Generated at: 2026-05-10T17:09:52Z
Summary
Source metadata CSV: /Users/huynhthanhdat/Workspace/iNexus/eKYB/ocr-captcha-finetuned/datasets/hf_dataset/source/metadata.csv
Rows total in source: 10954
Rows with label: 9000
Rows unlabeled: 1954
Rows skipped (missing image): 0
Exported labeled samples: 9000
Splits
train: 8100
validation: 900
test: 0
Files
HF imagefolder standard layout:
train/*.png… See the full description on the dataset page: https://huggingface.co/datasets/dathuynh1108/captcha-to-text.captcha_datasetcaptcha-synth-v3Deep_Captcha
DeepCaptcha Dataset: AI-Resistant CAPTCHA Benchmark
Overview
DeepCaptcha is a professional-grade dataset designed for benchmarking the AI-resistance and robustness of computer vision models. It contains CAPTCHA images generated with varying levels of adversarial protection, specifically engineered to thwart common automated recognition attacks (CNNs, OCR, etc.) while remaining human-readable.
Dataset Structure
The dataset is organized into 5 primary difficulty… See the full description on the dataset page: https://huggingface.co/datasets/Knight07/Deep_Captcha.advanced-math-captchas-v4Captcha_image
6位混合验证码数据集
数据集描述
内容:包含21000张6位验证码图片,字符组合涵盖0-9、A-Z、a-z共62种字符
生成方式:程序合成
标注格式:每张图片文件名即对应标签(如3aB9Zq.jpg的标签为3aB9Zq)
文件结构
dataset/
├── train/ # 训练集1万张
│ ├── 0aB1Cd.jpg
│ └── ...
├── test/ # 测试集1万张
├── val/ # 验证集1000张
captcha-assetscaptcha
