CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ed001 /ds-coder-instruct-v2 Dataset Card for DS Coder Instruct v2 Dataset Changes from v1: Added WizardLM evol data science samples Removed R samples from v2 DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2). The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.tabulartext-generation10K<n<100K13 likes230 downloads3y agoHugging Face02fasterinnerlooper /codereviewertabular100K<n<1M1 likes225 downloads3y agoHugging Face03JohnBeanerson /strudel-coder Claude Code session traces for JohnBeanerson/strudel-coder This dataset contains redacted Claude Code session traces collected while working on https://github.com/ultralazr/strudel-coder.git. The traces were exported with cc-share-hf and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each file in the repo root is a redacted Claude Code session in its native JSONL format (one entry per line). HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/JohnBeanerson/strudel-coder.tabulartext-generationn<1K0 likes164 downloads4mo agoHugging Face04CodeResearch /CodeJudge-Eval CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding? If our project helps you, please give us a star ⭐ on GitHub to support us. 🙏🙏 Introduction Recent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model's code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel… See the full description on the dataset page: https://huggingface.co/datasets/CodeResearch/CodeJudge-Eval.tabular10K<n<100K2 likes112 downloads2y agoHugging Face05nikitamounier /swe-bench-verified-raw-traces-qwen3-coder SWE-bench Verified raw mini-SWE-agent traces Raw mini-SWE-agent trajectories from 20250802_mini-v1.0.0_qwen3-coder-480b-a35b-instruct for SWE-bench Verified. The raw/easy split uses exactly the 194 instance IDs from parsaidp/SWE-bench_Verified_easy. That public dataset contains SWE-bench Verified questions and Kimi-generated answers; this dataset uses only its instance IDs. The trajectory contents here are local mini-SWE-agent/Qwen traces. Files data/full.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/nikitamounier/swe-bench-verified-raw-traces-qwen3-coder.tabular1K<n<10K1 likes77 downloads4mo agoHugging Face06prithivMLmods /Coder-Stat Coder-Stat Dataset Overview The Coder-Stat dataset is a collection of programming-related data, including problem IDs, programming languages, original statuses, and source code snippets. This dataset is designed to assist in the analysis of coding patterns, error types, and performance metrics. Dataset Details Modalities Tabular: The dataset is structured in a tabular format. Text: Contains text data, including source code snippets. Formats… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Coder-Stat.tabulartext-classification10K<n<100K3 likes76 downloads2y agoHugging Face07LLMTeamAkiyama /cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder データ件数: 269,863 平均トークン数: 11674 最大トークン数: 31,184 合計トークン数: 3,150,447,484 ファイル形式: JSONL ファイルサイズ: 不明 加工内容 synthetic_sftを使用 トークン処理が重たいので、文字数でフィルター seed_question < 6000 generation < 80000 thinkタグ除去 が中途半端なものを除外 トークナイズ処理(速度向上アップデート 繰り返し除去 tabularquestion-answering100K<n<1M0 likes70 downloads1y agoHugging Face08CodeResearch /Code-Evol-Instruct-OSS Code-Evol-Instruct-OSS Summary Code-Evol-Instruct-OSS is a dataset that was generated with Code Evol-Instruct by prompting open-souce LLMs, WizardLM-13B-v1.2 and WizardCoder-34B-Python. The underlying process is explained in the paper code-evol-instruct. This algorithm gave birth to famous open-souce code LLMs, WizardCoder-Family. Our approach We did not use any closed-source LLMs. Our seed dataset is sourced from self-instruct-starcoder. We leverage the… See the full description on the dataset page: https://huggingface.co/datasets/CodeResearch/Code-Evol-Instruct-OSS.tabular1K<n<10K6 likes65 downloads3y agoHugging Face09open-llm-leaderboard /LeroyDyer___Spydaz_Web_AI_AGI_R1_OmG_Coder-detailsgated Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_OmG_Coder Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_OmG_Coder The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_OmG_Coder-details.tabular10K<n<100K0 likes55 downloads2y agoHugging Face10open-llm-leaderboard /tomasmcm__sky-t1-coder-32b-flash-detailsgated Dataset Card for Evaluation run of tomasmcm/sky-t1-coder-32b-flash Dataset automatically created during the evaluation run of model tomasmcm/sky-t1-coder-32b-flash The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tomasmcm__sky-t1-coder-32b-flash-details.tabular10K<n<100K0 likes52 downloads2y agoHugging Face11open-llm-leaderboard /prithivMLmods__Viper-Coder-Hybrid-v1.3-detailsgated Dataset Card for Evaluation run of prithivMLmods/Viper-Coder-Hybrid-v1.3 Dataset automatically created during the evaluation run of model prithivMLmods/Viper-Coder-Hybrid-v1.3 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Viper-Coder-Hybrid-v1.3-details.tabular10K<n<100K0 likes48 downloads2y agoHugging Face12open-llm-leaderboard /prithivMLmods__Viper-Coder-v1.1-detailsgated Dataset Card for Evaluation run of prithivMLmods/Viper-Coder-v1.1 Dataset automatically created during the evaluation run of model prithivMLmods/Viper-Coder-v1.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Viper-Coder-v1.1-details.tabular10K<n<100K0 likes44 downloads2y agoHugging Face13open-llm-leaderboard /theo77186__Qwen2.5-Coder-7B-Instruct-20241106-detailsgated Dataset Card for Evaluation run of theo77186/Qwen2.5-Coder-7B-Instruct-20241106 Dataset automatically created during the evaluation run of model theo77186/Qwen2.5-Coder-7B-Instruct-20241106 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/theo77186__Qwen2.5-Coder-7B-Instruct-20241106-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face14parsaidp /swe-bench-verified-raw-traces-qwen3-coder SWE-bench Verified raw mini-SWE-agent traces Raw mini-SWE-agent trajectories from 20250802_mini-v1.0.0_qwen3-coder-480b-a35b-instruct for SWE-bench Verified. The raw/easy split uses exactly the 194 instance IDs from parsaidp/SWE-bench_Verified_easy. That public dataset contains SWE-bench Verified questions and Kimi-generated answers; this dataset uses only its instance IDs. The trajectory contents here are local mini-SWE-agent/Qwen traces. Files data/full.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/parsaidp/swe-bench-verified-raw-traces-qwen3-coder.tabular1K<n<10K0 likes41 downloads4mo agoHugging Face15open-llm-leaderboard /rombodawg__Rombos-Coder-V2.5-Qwen-14b-detailsgated Dataset Card for Evaluation run of rombodawg/Rombos-Coder-V2.5-Qwen-14b Dataset automatically created during the evaluation run of model rombodawg/Rombos-Coder-V2.5-Qwen-14b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rombodawg__Rombos-Coder-V2.5-Qwen-14b-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face16reallexi /lexi-coder-v3-datasest lexi-coder-v3-datasest lexi-coder-v3-datasest by Reallexi LLC AI Model Builder — llm.reallexi.io Copyright (c) 2026 Reallexi LLC. All rights reserved. A retrieval index built by Reallexi AI Model Builder: source text chunked, embedded, and stored for nearest-neighbor retrieval. This is not a causal-language-model checkpoint and cannot be loaded with AutoModelForCausalLM. Contents Source data lizn-zn/CodeNet_Extracted Documents indexed 1,000,000… See the full description on the dataset page: https://huggingface.co/datasets/reallexi/lexi-coder-v3-datasest.tabular1M<n<10M1 likes39 downloads2mo agoHugging Face17open-llm-leaderboard /yasserrmd__Coder-GRPO-3B-detailsgated Dataset Card for Evaluation run of yasserrmd/Coder-GRPO-3B Dataset automatically created during the evaluation run of model yasserrmd/Coder-GRPO-3B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yasserrmd__Coder-GRPO-3B-details.tabular10K<n<100K0 likes37 downloads2y agoHugging Face18open-llm-leaderboard /01-ai__Yi-Coder-9B-Chat-detailsgated Dataset Card for Evaluation run of 01-ai/Yi-Coder-9B-Chat Dataset automatically created during the evaluation run of model 01-ai/Yi-Coder-9B-Chat The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-Coder-9B-Chat-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face19open-llm-leaderboard /Qwen__Qwen2.5-Coder-7B-Instruct-detailsgated Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-7B-Instruct-details.tabular10K<n<100K0 likes35 downloads2y agoHugging Face20open-llm-leaderboard /Etherll__Qwen2.5-Coder-7B-Instruct-Ties-detailsgated Dataset Card for Evaluation run of Etherll/Qwen2.5-Coder-7B-Instruct-Ties Dataset automatically created during the evaluation run of model Etherll/Qwen2.5-Coder-7B-Instruct-Ties The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Etherll__Qwen2.5-Coder-7B-Instruct-Ties-details.tabular10K<n<100K0 likes35 downloads2y agoHugging Face21davidberenstein1957 /qwen2.5-coder-0.5b-openai_humaneval Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/qwen2.5-coder-0.5b-openai_humaneval.tabularn<1K0 likes34 downloads2y agoHugging Face22open-llm-leaderboard /LeroyDyer___Spydaz_Web_AI_AGI_R1_Student_Coder-detailsgated Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_Student_Coder Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_Student_Coder The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_Student_Coder-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face23open-llm-leaderboard /LeroyDyer___Spydaz_Web_AI_AGI_R1_Teacher_Coder-detailsgated Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_Teacher_Coder Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_Teacher_Coder The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_Teacher_Coder-details.tabular10K<n<100K0 likes31 downloads2y agoHugging Face24open-llm-leaderboard /Z1-Coder__Z1-Coder-7B-detailsgated Dataset Card for Evaluation run of Z1-Coder/Z1-Coder-7B Dataset automatically created during the evaluation run of model Z1-Coder/Z1-Coder-7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Z1-Coder__Z1-Coder-7B-details.tabular10K<n<100K0 likes29 downloads2y agoHugging Face25open-llm-leaderboard /TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Ins-Rule-detailsgated Dataset Card for Evaluation run of TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Ins-Rule Dataset automatically created during the evaluation run of model TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Ins-Rule The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Ins-Rule-details.tabular10K<n<100K0 likes28 downloads2y agoHugging Face26reallexi /lexi-coder-v4-dataset lexi-coder-v4-dataset lexi-coder-v4-dataset by Reallexi LLC AI Model Builder — llm.reallexi.io Copyright (c) 2026 Reallexi LLC. All rights reserved. A retrieval index built by Reallexi AI Model Builder: source text chunked, embedded, and stored for nearest-neighbor retrieval. This is not a causal-language-model checkpoint and cannot be loaded with AutoModelForCausalLM. Contents Source data [Nan-Do/code-search-net-python + google-research-datasets/mbpp… See the full description on the dataset page: https://huggingface.co/datasets/reallexi/lexi-coder-v4-dataset.tabular100K<n<1M0 likes28 downloads1mo agoHugging Face27open-llm-leaderboard /Qwen__Qwen2.5-Coder-14B-Instruct-detailsgated Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-14B-Instruct-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face28open-llm-leaderboard /Qwen__Qwen2.5-Coder-32B-Instruct-detailsgated Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-32B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-32B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-32B-Instruct-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face29open-llm-leaderboard /prithivMLmods__Viper-Coder-HybridMini-v1.3-detailsgated Dataset Card for Evaluation run of prithivMLmods/Viper-Coder-HybridMini-v1.3 Dataset automatically created during the evaluation run of model prithivMLmods/Viper-Coder-HybridMini-v1.3 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Viper-Coder-HybridMini-v1.3-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face30open-llm-leaderboard /TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Base-Rule-detailsgated Dataset Card for Evaluation run of TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Base-Rule Dataset automatically created during the evaluation run of model TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Base-Rule The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Base-Rule-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.