cyberco/CAIA_0927_Leaderboard
0
CAIA
CAIA is a benchmark for evaluating AI agents in adversarial, high-stakes financial environments.
For further details, please refer to the following resources:
- Paper: https://arxiv.org/pdf/2509.xxxxx
- Project Page: https://www.caiba.ai/
- Github: https://github.com/caiba-ai/caia-benchmark-0927
- CAIA Dataset: https://huggingface.co/datasets/cyberco/caia-0927
- Leaderboard: https://huggingface.co/spaces/cyberco/CAIA0927Leaderboard
- Point of Contact: James Dai
Citation @article{CAIA, title={When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets}, author={Dai, Zeshi and Peng, Zimo and Cheng, Zerui and Li, Yihe Ryan}, journal={arXiv preprint arXiv:2509.xxxxx}, year={2025} }
