CoolFace
Datasetpublic

m-a-p/AetherCode

AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions Introduction Competitive programming has emerged as a critical benchmark for evaluating the reasoning and coding capabilities of Large Language Models (LLMs). Despite impressive progress on existing benchmarks, we argue that current evaluations overstate model proficiency, masking a substantial gap between LLMs and elite human programmers. This gap… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/AetherCode.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
8likes688downloads
evaluation_result.png4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:a34c20c2a343f7fe1d3d691011319c623af4f762227711688ddaeb211874b3943size 2143204