datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
impossible_livecodebenchLiveCodeBench-Proexecution-v2livecodebench_code_generation_lite_parquet
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
🏠 Home Page •
💻 GitHub Repository •
🏆 Leaderboard •
📄 Paper
Change Log
Since LiveCodeBench is a continuously updated benchmark, we provide different versions of the dataset. Particularly, we provide the following versions of the dataset:
release_v1: The initial release of the dataset with problems released between May 2023 and Mar 2024 containing 400… See the full description on the dataset page: https://huggingface.co/datasets/marianna13/livecodebench_code_generation_lite_parquet.test_generation
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
🏠 Home Page •
💻 GitHub Repository •
🏆 Leaderboard •
LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs.
Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution.
This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/test_generation.livecodebench-code_generation_liteexecution
Dataset Card for "livecodebench-execute"
LiveCodeBench-v5livecodebench_code_generation_litelivecodebench-v6livecodebench_lite_filteredlivecodebench-code_generation_liteqwen3-8b-livecodebench-subset-v6-n128LiveCodeBench_V2_UTRL_evalTest set for evaluating LLM-based unit test generation capabilities, built upon LiveCodeBench-v2.
problem_statement: Description of the programming problem in LivecCodeBench-v2.
gt_test_cases: Ground-truth test cases to evaluate the correctness of the arbitrary code solutions.
sampled_code: 64 code solutions sampled from Qwen3-4B and Qwen3-8B.
Following evaluation scheme in Lee et al., 2026, Unit test generated by LLMs can be evaluated by the following metrics:
Best-of-N improvement:… See the full description on the dataset page: https://huggingface.co/datasets/dgjun32/LiveCodeBench_V2_UTRL_eval.qwen3-8b-livecodebench-v6-n128Ag-LiveCodeBench-XThis repository contains the multi-PL variant of LiveCodeBench, prepared in the Agnostics project.
Find out more about the dataset and the related artifacts on the project website.
The easiest way to benchmark a model on this dataset is with our scripts.
livecodebench
LiveCodeBench for Code-LLaVA
This dataset contains the LiveCodeBench code generation benchmark prepared for
Code-LLaVA evaluation.
Source
Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Repository: https://github.com/LiveCodeBench/LiveCodeBench
Version: release_v6 (May 2023 - Apr 2025, 1055 problems)
Dataset Structure
Two configurations are available:
memwrap: Problems with <|memory_start|> /… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/livecodebench.livecodebench-merging-leaderboard
LiveCodeBench v6 Evaluation Leaderboard
Evaluation results for cross-capability merging of OLMo-3 and OLMo-3.1 RL-Zero models on 454 coding problems.
Evaluation
We followed the evaluation guidelines and prompts from OLMo 3. Best effort was made to ensure reported numbers are as accurate as possible.
Code: pmahdavi/modal-eval
Leaderboard
Model
pass@4
pass@1
Loop Rate
Qwen/Qwen3-4B-Thinking-2507
54.6%
45.4%
0.4%
pmahdavi/Olmo-3-7B-Think-Math-Code… See the full description on the dataset page: https://huggingface.co/datasets/pmahdavi/livecodebench-merging-leaderboard.livecodebench_lite_contaminatedlivecodebench-execute-v2r1_livecodebenchlivecodebenchlivecodebench_pro_cppqwen3-14b-livecodebench-subset-v6-n128B_livecodebench_lite_v3livecodebench_unit_test_error_240_sampleslivecodebench-subset-v6livecodebench_v6_rawUnroll-Qwen2.5-7B-Instruct_1754646934_eval_6419_livecodebench_skip_ffn_idx_8_v2
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754646934_eval_6419_livecodebench_skip_ffn_idx_8_v2
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 20.38% ± 0.49%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
21.33%
109
511
2
20.35%
104
511
3
20.55%
105
511
4
18.59%
95
511
5
21.92%
112
511
6
19.57%
100
511
livecodebench_lite_v2_testbank_re
