m-a-p/AetherCode
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions Introduction Competitive programming has emerged as a critical benchmark for evaluating the reasoning and coding capabilities of Large Language Models (LLMs). Despite impressive progress on existing benchmarks, we argue that current evaluations overstate model proficiency, masking a substantial gap between LLMs and elite human programmers. This gap… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/AetherCode.
Improve dataset card: Add metadata, update Hugging Face and paper badges (#2)
Update README.md
Update README.md
Update README.md
Update README.md
Create DISCLAIMER
Update README.md
Upload 2 files
Update README.md
Update README.md
Create LICENSE
Update README.md
Update README.md
Upload dataset
Upload dataset
Delete README.md
Delete v1_2025
Upload dataset
initial commit
