datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CAP-Bench
CAP-Bench
A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception.
CAP-Bench evaluates browser agents on Cross-site workflows, complex Actions, and challenging visual Perception. The full benchmark contains 420 tasks across 108 real-world websites in 24 functional domains. Each task requires on average 7 complex execution operations and 4 perception challenges, substantially exceeding the difficulty of prior browser-agent benchmarks.
This… See the full description on the dataset page: https://huggingface.co/datasets/Warrior0302/CAP-Bench.captum-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: captum
Documentation Data Source Link: https://captum.ai/docs/introduction
Data Source License: https://github.com/pytorch/captum?tab=BSD-3-Clause-1-ov-file#readme
Data Source Authors: Observable AI Benchmarks by Data Agents © 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
captum-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: captum
Documentation Data Source Link: https://captum.ai/docs/introduction
Data Source License: https://github.com/pytorch/captum?tab=BSD-3-Clause-1-ov-file#readme
Data Source Authors: Observable AI Benchmarks by Data Agents © 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
