datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Uncensored-CodeLlama
Dataset Card for Evaluation run of ehartford/WizardLM-1.0-Uncensored-CodeLlama-34b
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-1.0-Uncensored-CodeLlama-34b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/akpsahan/Uncensored-CodeLlama.cortex-codellamaCodeLlama-2-20k
CodeLlama-2-20k: A Llama 2 Version of CodeAlpaca
This dataset is the sahil2801/CodeAlpaca-20k dataset with the Llama 2 prompt format described here.
Here is the code I used to format it:
from datasets import load_dataset
# Load the dataset
dataset = load_dataset('sahil2801/CodeAlpaca-20k')
# Define a function to merge the three columns into one
def merge_columns(example):
if example['input']:
merged = f"<s>[INST] <<SYS>>\nBelow is an instruction that describes a task… See the full description on the dataset page: https://huggingface.co/datasets/mlabonne/CodeLlama-2-20k.Magicoder-Evol-Instruct-500-CodeLlama-70b-tokenized-0.5-Special-Tokenhl-codellama-chat-response
Dataset Card for "hl-codellama-chat-response"
More Information needed
vuejs-nuxt-tailwind-codellama-examplesMagicoder-Evol-Instruct-5000-CodeLlama-70b-tokenized-0.5-v2spider-codefuse-codellama-34b-temp1.0codellama_unity3d_v2spider-codellama-70b-instruct-hf-temp0.0-pandasvuejs-codellamaspider-codellama-13b-instruct-hf-temp0.0-pandashl-codellama-chat-response-v2
Dataset Card for "hl-codellama-chat-response-v2"
More Information needed
spider-codellama-7b-instruct-hf-temp0.0-pandascodellama-finetuningSecCoderX_CodeLlama_7b_GRPO_dataset
Citation
If you find our work helpful, feel free to give us a cite.
@misc{wu2026securecodegenerationonline,
title={Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model},
author={Tianyi Wu and Mingzhe Du and Yue Liu and Chengran Yang and Terry Yue Zhuo and Jiaheng Zhang and See-Kiong Ng},
year={2026},
eprint={2602.07422},
archivePrefix={arXiv},
primaryClass={cs.CR},
url={https://arxiv.org/abs/2602.07422}… See the full description on the dataset page: https://huggingface.co/datasets/SecCoderX/SecCoderX_CodeLlama_7b_GRPO_dataset.cortex-codellama-fchumaneval-cot-codellama70bMagicoder-Evol-Instruct-10000-CodeLlama-70b-tokenized-0.5-v2own_target_data_codellamavuejs-nuxt-tailwind-codellamaCode-Llama-Customcodellama_unity3d_test_50code_llama3.1_405b_instruct_10000wikisql_codellama_1000
Dataset Card for "wikisql_codellama_1000"
More Information needed
Magicoder-Evol-Instruct-250-CodeLlama-70b-tokenized-0.5-Special-TokenMagicoder-Evol-Instruct-1000-CodeLlama-70b-tokenized-0.5-Special-Tokencode-llama-13b-fitm-mask-heldout-1-epoch-base-dataAtlas_CodeLlama7bInstruct_Tokenizedcodellama_unity3d_v1
