CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fjzzq2002 /impossible_livecodebenchtextn<1K1 likes1.9k downloads1y agoHugging Face02QAQAQAQAQ /LiveCodeBench-Progatedtabular1K<n<10K10 likes1.3k downloads11mo agoHugging Face03livecodebench /execution-v2tabularn<1K5 likes1.1k downloads2y agoHugging Face04marianna13 /livecodebench_code_generation_lite_parquet LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • 📄 Paper Change Log Since LiveCodeBench is a continuously updated benchmark, we provide different versions of the dataset. Particularly, we provide the following versions of the dataset: release_v1: The initial release of the dataset with problems released between May 2023 and Mar 2024 containing 400… See the full description on the dataset page: https://huggingface.co/datasets/marianna13/livecodebench_code_generation_lite_parquet.text10K<n<100K0 likes835 downloads9mo agoHugging Face05livecodebench /test_generation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs. Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution. This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/test_generation.textn<1K8 likes560 downloads2y agoHugging Face06namanbnsl /livecodebench-code_generation_litetext1K<n<10K2 likes556 downloads5mo agoHugging Face07livecodebench /execution Dataset Card for "livecodebench-execute" textn<1K6 likes352 downloads3y agoHugging Face08PrimeIntellect /LiveCodeBench-v5textn<1K0 likes333 downloads1y agoHugging Face09RewardGuided /livecodebench_code_generation_litetext1K<n<10K0 likes146 downloads20d agoHugging Face10drproduck /livecodebench-v6text1K<n<10K0 likes145 downloads1y agoHugging Face11cassanof /livecodebench_lite_filteredtextn<1K0 likes121 downloads2y agoHugging Face12ali-elganzory /livecodebench-code_generation_litetext1K<n<10K0 likes116 downloads6mo agoHugging Face13drproduck /qwen3-8b-livecodebench-subset-v6-n128textn<1K0 likes103 downloads1y agoHugging Face14dgjun32 /LiveCodeBench_V2_UTRL_evalTest set for evaluating LLM-based unit test generation capabilities, built upon LiveCodeBench-v2. problem_statement: Description of the programming problem in LivecCodeBench-v2. gt_test_cases: Ground-truth test cases to evaluate the correctness of the arbitrary code solutions. sampled_code: 64 code solutions sampled from Qwen3-4B and Qwen3-8B. Following evaluation scheme in Lee et al., 2026, Unit test generated by LLMs can be evaluated by the following metrics: Best-of-N improvement:… See the full description on the dataset page: https://huggingface.co/datasets/dgjun32/LiveCodeBench_V2_UTRL_eval.textn<1K0 likes86 downloads8mo agoHugging Face15drproduck /qwen3-8b-livecodebench-v6-n128textn<1K0 likes70 downloads1y agoHugging Face16nuprl /Ag-LiveCodeBench-XThis repository contains the multi-PL variant of LiveCodeBench, prepared in the Agnostics project. Find out more about the dataset and the related artifacts on the project website. The easiest way to benchmark a model on this dataset is with our scripts. textn<1K1 likes62 downloads1y agoHugging Face17tonychenxyz /livecodebench LiveCodeBench for Code-LLaVA This dataset contains the LiveCodeBench code generation benchmark prepared for Code-LLaVA evaluation. Source Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Repository: https://github.com/LiveCodeBench/LiveCodeBench Version: release_v6 (May 2023 - Apr 2025, 1055 problems) Dataset Structure Two configurations are available: memwrap: Problems with <|memory_start|> /… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/livecodebench.texttext-generation1K<n<10K0 likes53 downloads8mo agoHugging Face18pmahdavi /livecodebench-merging-leaderboard LiveCodeBench v6 Evaluation Leaderboard Evaluation results for cross-capability merging of OLMo-3 and OLMo-3.1 RL-Zero models on 454 coding problems. Evaluation We followed the evaluation guidelines and prompts from OLMo 3. Best effort was made to ensure reported numbers are as accurate as possible. Code: pmahdavi/modal-eval Leaderboard Model pass@4 pass@1 Loop Rate Qwen/Qwen3-4B-Thinking-2507 54.6% 45.4% 0.4% pmahdavi/Olmo-3-7B-Think-Math-Code… See the full description on the dataset page: https://huggingface.co/datasets/pmahdavi/livecodebench-merging-leaderboard.tabulartext-generation10K<n<100K0 likes52 downloads8mo agoHugging Face19cassanof /livecodebench_lite_contaminatedtextn<1K0 likes48 downloads2y agoHugging Face20minimario /livecodebench-execute-v2text1K<n<10K1 likes47 downloads3y agoHugging Face21YangZhoumill /r1_livecodebenchtextn<1K0 likes45 downloads4mo agoHugging Face22DongfuJiang /livecodebenchtextn<1K0 likes42 downloads1y agoHugging Face23llamastack /livecodebench_pro_cpptextn<1K0 likes42 downloads6mo agoHugging Face24drproduck /qwen3-14b-livecodebench-subset-v6-n128textn<1K0 likes39 downloads1y agoHugging Face25codegenning /B_livecodebench_lite_v3textn<1K0 likes38 downloads2y agoHugging Face26rmcc11 /livecodebench_unit_test_error_240_samplestextn<1K0 likes37 downloads1y agoHugging Face27drproduck /livecodebench-subset-v6textn<1K0 likes33 downloads1y agoHugging Face28MikeZheng777 /livecodebench_v6_rawtext1K<n<10K0 likes33 downloads2mo agoHugging Face29chengfu0118 /Unroll-Qwen2.5-7B-Instruct_1754646934_eval_6419_livecodebench_skip_ffn_idx_8_v2 chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754646934_eval_6419_livecodebench_skip_ffn_idx_8_v2 Precomputed model outputs for evaluation. Evaluation Results LiveCodeBench Average Accuracy: 20.38% ± 0.49% Number of Runs: 6 Run Accuracy Questions Solved Total Questions 1 21.33% 109 511 2 20.35% 104 511 3 20.55% 105 511 4 18.59% 95 511 5 21.92% 112 511 6 19.57% 100 511 tabular1K<n<10K0 likes31 downloads1y agoHugging Face30codegenning /livecodebench_lite_v2_testbank_retextn<1K0 likes30 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.