CoolFace
Datasetpublic

berryccc1/EVM-QuestBench

EVM-QuestBench EVM-QuestBench is an execution-grounded benchmark for evaluating whether large language models and AI agents can translate natural-language blockchain intent into executable transaction code that produces the intended EVM state transition. Released with the ACL 2026 Long Paper by Pei Yang, Wanyi Chen, Ke Wang, Lynn Ai, Eric Yang, and Tianyu Shi, the benchmark contains 107 expert-authored tasks: 62 atomic tasks and 45 composite workflows. Unlike code-similarity… See the full description on the dataset page: https://huggingface.co/datasets/berryccc1/EVM-QuestBench.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes29downloads

berryccc1/EVM-QuestBench · main · files are served by the source, never re-hosted here