datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ds-coder-instruct-v2
Dataset Card for DS Coder Instruct v2 Dataset
Changes from v1:
Added WizardLM evol data science samples
Removed R samples from v2
DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on
data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2).
The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.codereviewerstrudel-coder
Claude Code session traces for JohnBeanerson/strudel-coder
This dataset contains redacted Claude Code session traces collected while working on https://github.com/ultralazr/strudel-coder.git. The traces were exported with cc-share-hf and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each file in the repo root is a redacted Claude Code session in its native JSONL format (one entry per line). HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/JohnBeanerson/strudel-coder.CodeJudge-Eval
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
If our project helps you, please give us a star ⭐ on GitHub to support us. 🙏🙏
Introduction
Recent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model's code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel… See the full description on the dataset page: https://huggingface.co/datasets/CodeResearch/CodeJudge-Eval.swe-bench-verified-raw-traces-qwen3-coder
SWE-bench Verified raw mini-SWE-agent traces
Raw mini-SWE-agent trajectories from 20250802_mini-v1.0.0_qwen3-coder-480b-a35b-instruct for SWE-bench Verified.
The raw/easy split uses exactly the 194 instance IDs from
parsaidp/SWE-bench_Verified_easy. That public dataset contains SWE-bench
Verified questions and Kimi-generated answers; this dataset uses only its
instance IDs. The trajectory contents here are local mini-SWE-agent/Qwen traces.
Files
data/full.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/nikitamounier/swe-bench-verified-raw-traces-qwen3-coder.Coder-Stat
Coder-Stat Dataset
Overview
The Coder-Stat dataset is a collection of programming-related data, including problem IDs, programming languages, original statuses, and source code snippets. This dataset is designed to assist in the analysis of coding patterns, error types, and performance metrics.
Dataset Details
Modalities
Tabular: The dataset is structured in a tabular format.
Text: Contains text data, including source code snippets.
Formats… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Coder-Stat.cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder
データ件数: 269,863
平均トークン数: 11674
最大トークン数: 31,184
合計トークン数: 3,150,447,484
ファイル形式: JSONL
ファイルサイズ: 不明
加工内容
synthetic_sftを使用
トークン処理が重たいので、文字数でフィルター
seed_question < 6000
generation < 80000
thinkタグ除去 が中途半端なものを除外
トークナイズ処理(速度向上アップデート
繰り返し除去
Code-Evol-Instruct-OSS
Code-Evol-Instruct-OSS
Summary
Code-Evol-Instruct-OSS is a dataset that was generated with Code Evol-Instruct by prompting open-souce LLMs, WizardLM-13B-v1.2 and WizardCoder-34B-Python.
The underlying process is explained in the paper code-evol-instruct. This algorithm gave birth to famous open-souce code LLMs, WizardCoder-Family.
Our approach
We did not use any closed-source LLMs.
Our seed dataset is sourced from self-instruct-starcoder.
We leverage the… See the full description on the dataset page: https://huggingface.co/datasets/CodeResearch/Code-Evol-Instruct-OSS.LeroyDyer___Spydaz_Web_AI_AGI_R1_OmG_Coder-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_OmG_Coder
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_OmG_Coder
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_OmG_Coder-details.tomasmcm__sky-t1-coder-32b-flash-details
Dataset Card for Evaluation run of tomasmcm/sky-t1-coder-32b-flash
Dataset automatically created during the evaluation run of model tomasmcm/sky-t1-coder-32b-flash
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tomasmcm__sky-t1-coder-32b-flash-details.prithivMLmods__Viper-Coder-Hybrid-v1.3-details
Dataset Card for Evaluation run of prithivMLmods/Viper-Coder-Hybrid-v1.3
Dataset automatically created during the evaluation run of model prithivMLmods/Viper-Coder-Hybrid-v1.3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Viper-Coder-Hybrid-v1.3-details.prithivMLmods__Viper-Coder-v1.1-details
Dataset Card for Evaluation run of prithivMLmods/Viper-Coder-v1.1
Dataset automatically created during the evaluation run of model prithivMLmods/Viper-Coder-v1.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Viper-Coder-v1.1-details.theo77186__Qwen2.5-Coder-7B-Instruct-20241106-details
Dataset Card for Evaluation run of theo77186/Qwen2.5-Coder-7B-Instruct-20241106
Dataset automatically created during the evaluation run of model theo77186/Qwen2.5-Coder-7B-Instruct-20241106
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/theo77186__Qwen2.5-Coder-7B-Instruct-20241106-details.swe-bench-verified-raw-traces-qwen3-coder
SWE-bench Verified raw mini-SWE-agent traces
Raw mini-SWE-agent trajectories from 20250802_mini-v1.0.0_qwen3-coder-480b-a35b-instruct for SWE-bench Verified.
The raw/easy split uses exactly the 194 instance IDs from parsaidp/SWE-bench_Verified_easy. That
public dataset contains SWE-bench Verified questions and Kimi-generated answers; this
dataset uses only its instance IDs. The trajectory contents here are local
mini-SWE-agent/Qwen traces.
Files
data/full.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/parsaidp/swe-bench-verified-raw-traces-qwen3-coder.rombodawg__Rombos-Coder-V2.5-Qwen-14b-details
Dataset Card for Evaluation run of rombodawg/Rombos-Coder-V2.5-Qwen-14b
Dataset automatically created during the evaluation run of model rombodawg/Rombos-Coder-V2.5-Qwen-14b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rombodawg__Rombos-Coder-V2.5-Qwen-14b-details.lexi-coder-v3-datasest
lexi-coder-v3-datasest
lexi-coder-v3-datasest by Reallexi LLC AI Model Builder — llm.reallexi.io
Copyright (c) 2026 Reallexi LLC. All rights reserved.
A retrieval index built by Reallexi AI Model Builder: source text chunked, embedded, and stored for nearest-neighbor retrieval. This is not a causal-language-model checkpoint and cannot be loaded with AutoModelForCausalLM.
Contents
Source data
lizn-zn/CodeNet_Extracted
Documents indexed
1,000,000… See the full description on the dataset page: https://huggingface.co/datasets/reallexi/lexi-coder-v3-datasest.yasserrmd__Coder-GRPO-3B-details
Dataset Card for Evaluation run of yasserrmd/Coder-GRPO-3B
Dataset automatically created during the evaluation run of model yasserrmd/Coder-GRPO-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yasserrmd__Coder-GRPO-3B-details.01-ai__Yi-Coder-9B-Chat-details
Dataset Card for Evaluation run of 01-ai/Yi-Coder-9B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-Coder-9B-Chat
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-Coder-9B-Chat-details.Qwen__Qwen2.5-Coder-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-7B-Instruct-details.Etherll__Qwen2.5-Coder-7B-Instruct-Ties-details
Dataset Card for Evaluation run of Etherll/Qwen2.5-Coder-7B-Instruct-Ties
Dataset automatically created during the evaluation run of model Etherll/Qwen2.5-Coder-7B-Instruct-Ties
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Etherll__Qwen2.5-Coder-7B-Instruct-Ties-details.qwen2.5-coder-0.5b-openai_humaneval
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/qwen2.5-coder-0.5b-openai_humaneval.LeroyDyer___Spydaz_Web_AI_AGI_R1_Student_Coder-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_Student_Coder
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_Student_Coder
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_Student_Coder-details.LeroyDyer___Spydaz_Web_AI_AGI_R1_Teacher_Coder-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_Teacher_Coder
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_Teacher_Coder
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_Teacher_Coder-details.Z1-Coder__Z1-Coder-7B-details
Dataset Card for Evaluation run of Z1-Coder/Z1-Coder-7B
Dataset automatically created during the evaluation run of model Z1-Coder/Z1-Coder-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Z1-Coder__Z1-Coder-7B-details.TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Ins-Rule-details
Dataset Card for Evaluation run of TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Ins-Rule
Dataset automatically created during the evaluation run of model TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Ins-Rule
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Ins-Rule-details.lexi-coder-v4-dataset
lexi-coder-v4-dataset
lexi-coder-v4-dataset by Reallexi LLC AI Model Builder — llm.reallexi.io
Copyright (c) 2026 Reallexi LLC. All rights reserved.
A retrieval index built by Reallexi AI Model Builder: source text chunked, embedded, and stored for nearest-neighbor retrieval. This is not a causal-language-model checkpoint and cannot be loaded with AutoModelForCausalLM.
Contents
Source data
[Nan-Do/code-search-net-python + google-research-datasets/mbpp… See the full description on the dataset page: https://huggingface.co/datasets/reallexi/lexi-coder-v4-dataset.Qwen__Qwen2.5-Coder-14B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-14B-Instruct-details.Qwen__Qwen2.5-Coder-32B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-32B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-32B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-32B-Instruct-details.prithivMLmods__Viper-Coder-HybridMini-v1.3-details
Dataset Card for Evaluation run of prithivMLmods/Viper-Coder-HybridMini-v1.3
Dataset automatically created during the evaluation run of model prithivMLmods/Viper-Coder-HybridMini-v1.3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Viper-Coder-HybridMini-v1.3-details.TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Base-Rule-details
Dataset Card for Evaluation run of TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Base-Rule
Dataset automatically created during the evaluation run of model TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Base-Rule
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Base-Rule-details.
