CoolFace
Datasetpublic

m-a-p/AetherCode

AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions Introduction Competitive programming has emerged as a critical benchmark for evaluating the reasoning and coding capabilities of Large Language Models (LLMs). Despite impressive progress on existing benchmarks, we argue that current evaluations overstate model proficiency, masking a substantial gap between LLMs and elite human programmers. This gap… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/AetherCode.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
8likes659downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

m-a-p/AetherCode · main · files are served by the source, never re-hosted here