datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magpie-Qwen2.5-Coder-Pro-300K-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Coder-Pro-300K-v0.1.details_Qwen__Qwen2.5-Coder-14B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-14B-Instruct.details_Qwen__Qwen2.5-Coder-7B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-7B-Instruct.Qwen2.5-Coder-0.5B-Flutter-steps-eval
Qwen2.5-Coder-0.5B Flutter — Steps Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-steps,
a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In steps
mode, the model is given an existing file and an edit instruction and generates a
sequence of localized search/replace edit actions, each mechanically applied to the
current file state before the next action is generated, until the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-steps-eval.Qwen2.5-Coder-0.5B-Flutter-direct-eval
Qwen2.5-Coder-0.5B Flutter — Direct Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct,
a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In direct
mode, the model is given an existing file and an edit instruction and generates the
complete modified file in a single forward pass (as opposed to the steps /
iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct-eval.all-do-Qwen2.5-Coder-32B-Instructdetails_Qwen__Qwen2.5-Coder-14B
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-14B.theo77186__Qwen2.5-Coder-7B-Instruct-20241106-details
Dataset Card for Evaluation run of theo77186/Qwen2.5-Coder-7B-Instruct-20241106
Dataset automatically created during the evaluation run of model theo77186/Qwen2.5-Coder-7B-Instruct-20241106
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/theo77186__Qwen2.5-Coder-7B-Instruct-20241106-details.details_Qwen__Qwen2.5-Coder-3B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-3B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-3B-Instruct.
The dataset is composed of 1 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/details_Qwen__Qwen2.5-Coder-3B-Instruct.Qwen__Qwen2.5-Coder-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-7B-Instruct-details.Etherll__Qwen2.5-Coder-7B-Instruct-Ties-details
Dataset Card for Evaluation run of Etherll/Qwen2.5-Coder-7B-Instruct-Ties
Dataset automatically created during the evaluation run of model Etherll/Qwen2.5-Coder-7B-Instruct-Ties
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Etherll__Qwen2.5-Coder-7B-Instruct-Ties-details.qwen2.5-coder-0.5b-openai_humaneval
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/qwen2.5-coder-0.5b-openai_humaneval.DeltaSecommits_qwen2.5-coder-14b-instruct_tokenized_v2_vulnerableru_codefeedback_python_Qwen2.5-Coder-32B-Instruct-GPTQ-Int8_sample
ru_Code-Feedback
Вопросы python Code-Feedback
Решение и unit-test с результатами python исполнения.
Made with Qwen2.5-Coder-32B-Instruct-GPTQ-Int8
ru_eval_status
count
OK
2554
Exception
2337
SyntaxError
518
Timeout
79
ioi-eval-sglang_Qwen_Qwen2.5-Coder-32B-Instruct-prompt-mem-limit-fixhumaneval___2txt___Qwen2.5-Math-7B-Instruct___Qwen2.5-Coder-7B-Instruct___from_evalplusTIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Ins-Rule-details
Dataset Card for Evaluation run of TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Ins-Rule
Dataset automatically created during the evaluation run of model TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Ins-Rule
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Ins-Rule-details.gsm8k___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_InstructQwen__Qwen2.5-Coder-14B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-14B-Instruct-details.Qwen__Qwen2.5-Coder-32B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-32B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-32B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-32B-Instruct-details.Magpie-Qwen2.5-Coder-Pro-300K-Query-Positive-PairTIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Base-Rule-details
Dataset Card for Evaluation run of TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Base-Rule
Dataset automatically created during the evaluation run of model TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Base-Rule
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Base-Rule-details.qwen2.5-coder-7b-inst-sleeper-agent-2024-completionsQwen__Qwen2.5-Coder-14B-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-14B-details.code_mbpp_Qwen2.5-Coder-0.5B-Instruct_temp0.1_num8_tests_mbpp_qwen-7b-easy_t0.0_n1SecCoderX_Qwen2.5_Coder_7B_GRPO_dataset
Citation
If you find our work helpful, feel free to give us a cite.
@misc{wu2026securecodegenerationonline,
title={Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model},
author={Tianyi Wu and Mingzhe Du and Yue Liu and Chengran Yang and Terry Yue Zhuo and Jiaheng Zhang and See-Kiong Ng},
year={2026},
eprint={2602.07422},
archivePrefix={arXiv},
primaryClass={cs.CR},
url={https://arxiv.org/abs/2602.07422}… See the full description on the dataset page: https://huggingface.co/datasets/SecCoderX/SecCoderX_Qwen2.5_Coder_7B_GRPO_dataset.DCAgent2_terminal_bench_2_Qwen_Qwen2.5-Coder-32B-Instruct-tracesQwen__Qwen2.5-Coder-32B-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-32B
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-32B-details.mbpp_Qwen2.5-Coder-14B-Instruct_t1.0_n8_generated_codehumaneval___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instruct
