Medli1108/Open_CaptchaWorld
Open CaptchaWorld Dataset This dataset accompanies the paper Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents. It contains 20 distinct CAPTCHA types, each testing different visual reasoning capabilities. The dataset is designed for evaluating the visual reasoning and interaction capabilities of Multimodal Large Language Model (MLLM)-powered agents. Project Page | Github The dataset includes: 20 CAPTCHA Types: A diverse set… See the full description on the dataset page: https://huggingface.co/datasets/Medli1108/Open_CaptchaWorld.
Open CaptchaWorld Dataset
This dataset accompanies the paper Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents. It contains 20 distinct CAPTCHA types, each testing different visual reasoning capabilities. The dataset is designed for evaluating the visual reasoning and interaction capabilities of Multimodal Large Language Model (MLLM)-powered agents.
The dataset includes:
- 20 CAPTCHA Types: A diverse set of visual puzzles testing various capabilities. See the Github repository for a full list.
- Web Interface: A clean, intuitive interface for human or AI interaction.
- API Endpoints: Programmatic access to puzzles and verification.
This dataset is useful for benchmarking and improving multimodal AI agents' performance on CAPTCHA-like challenges, a crucial step in deploying web agents for real-world tasks. The data is structured for easy integration into research and development pipelines.
