datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Uncensored-CodeLlama
Dataset Card for Evaluation run of ehartford/WizardLM-1.0-Uncensored-CodeLlama-34b
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-1.0-Uncensored-CodeLlama-34b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/akpsahan/Uncensored-CodeLlama.cortex-codellamaCodeLlama-2-20k
CodeLlama-2-20k: A Llama 2 Version of CodeAlpaca
This dataset is the sahil2801/CodeAlpaca-20k dataset with the Llama 2 prompt format described here.
Here is the code I used to format it:
from datasets import load_dataset
# Load the dataset
dataset = load_dataset('sahil2801/CodeAlpaca-20k')
# Define a function to merge the three columns into one
def merge_columns(example):
if example['input']:
merged = f"<s>[INST] <<SYS>>\nBelow is an instruction that describes a task… See the full description on the dataset page: https://huggingface.co/datasets/mlabonne/CodeLlama-2-20k.Pangpuriye-generated_by_LLama3-codeLlama
🤖 Super AI Engineer Development Program Season 4 - Pangpuriye House - Generated by LLama3+codeLlama
Pangpuriye's House Dataset - Generated Dataset from LLama3+codeLlama
The dataset is a pack of text generation from LLama3 and codeLlama. The dataset is set under cc-by-nc-2.0 license.
Content
The dataset consists of 54,033 rows of input, instruction, and output. Most of the context in the dataset is in Thai. Whereas, the output is generally the answers regarding… See the full description on the dataset page: https://huggingface.co/datasets/AIAT/Pangpuriye-generated_by_LLama3-codeLlama.codellama_unity3d_v2Magicoder-Evol-Instruct-500-CodeLlama-70b-tokenized-0.5-Special-Tokencodellama_java_pythoncodellama-threejshl-codellama-chat-response
Dataset Card for "hl-codellama-chat-response"
More Information needed
Magicoder-Evol-Instruct-5000-CodeLlama-70b-tokenized-0.5-v2spider-codefuse-codellama-34b-temp1.0vuejs-nuxt-tailwind-codellama-examplesAdaDecode-CodeLlama-34B-Instruct-GSM8Kcodellama-step-2spider-codellama-13b-instruct-hf-temp0.0-pandasspider-codellama-70b-instruct-hf-temp0.0-pandasAdaDecode-CodeLlama-34B-Instruct-HumanEvalvuejs-codellamaspider-codellama-7b-instruct-hf-temp0.0-pandashl-codellama-chat-response-v2
Dataset Card for "hl-codellama-chat-response-v2"
More Information needed
CodeLlama-FormatSecCoderX_CodeLlama_7b_GRPO_dataset
Citation
If you find our work helpful, feel free to give us a cite.
@misc{wu2026securecodegenerationonline,
title={Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model},
author={Tianyi Wu and Mingzhe Du and Yue Liu and Chengran Yang and Terry Yue Zhuo and Jiaheng Zhang and See-Kiong Ng},
year={2026},
eprint={2602.07422},
archivePrefix={arXiv},
primaryClass={cs.CR},
url={https://arxiv.org/abs/2602.07422}… See the full description on the dataset page: https://huggingface.co/datasets/SecCoderX/SecCoderX_CodeLlama_7b_GRPO_dataset.codellama-finetuningcortex-codellama-fcAdaDecode-CodeLlama-13B-Instruct-GSM8Kcmg-codellama
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/zhaospei/cmg-codellama.own_target_data_codellamaMagicoder-Evol-Instruct-10000-CodeLlama-70b-tokenized-0.5-v2humaneval-cot-codellama70bvuejs-nuxt-tailwind-codellama
