jie-jw-wu/HumanEvalComm
HumanEvalComm: Benchmarking the Communication Skills of Code Generation for LLMs and LLM Agent π Paper β’ π» GitHub Repository β’ π€ Dataset Viewer Dataset Description HumanEvalComm is a benchmark dataset for evaluating the communication skills of Large Language Models (LLMs) in code generation tasks. It is built upon the widely used HumanEval benchmark. HumanEvalComm contains 762 modified problem descriptions based on the 164 problems inβ¦ See the full description on the dataset page: https://huggingface.co/datasets/jie-jw-wu/HumanEvalComm.
0178
