CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nickrosh /Evol-Instruct-Code-80k-v1Open Source Implementation of Evol-Instruct-Code as described in the WizardCoder Paper. Code for the intruction generation can be found on Github as Evol-Teacher. text10K<n<100K251 likes13k downloads3y agoHugging Face02likaixin /InstructCoder Paper | Code | Blog InstructCoder (CodeInstruct): Empowering Language Models to Edit Code Updates May 23, 2023: Paper, code and data released. Overview InstructCoder is the first dataset designed to adapt LLMs for general code editing. It consists of 114,239 instruction-input-output triplets and covers multiple distinct code editing scenarios, generated by ChatGPT. LLaMA-33B finetuned on InstructCoder performs on par with ChatGPT on a… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/InstructCoder.texttext-generation100K<n<1M17 likes7.7k downloads2y agoHugging Face03BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.4k downloads9mo agoHugging Face04codeparrot /self-instruct-starcoder Self-instruct-starcoder Summary Self-instruct-starcoder is a dataset that was generated by prompting starcoder to generate new instructions based on some human-written seed instructions. The underlying process is explained in the paper self-instruct. This algorithm gave birth to famous machine generated datasets such as Alpaca and Code Alpaca which are two datasets obtained by prompting OpenAI text-davinci-003 engine. Our approach While our method is… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/self-instruct-starcoder.text1K<n<10K64 likes690 downloads3y agoHugging Face05TQRG /bigcodebench_codellama_codellama-7b-instruct-hf_tokenized1K<n<10K0 likes613 downloads9mo agoHugging Face06FineEnvs /repo2rlenv-code-instruct repo2rlenv-code-instruct Generated by Repo2RLEnv — turning real GitHub repositories into verifiable RL environments. 💡 Browse this dataset in your browser — click the badge above or open HuggingFaceH4/harbor-visualiser to inspect every task's spec, instruction, oracle patch, test script, and Dockerfile. Source repos (5): encode/starlette pallets/click pallets/flask psf/requests python-attrs/attrs Pipeline: code_instruct Tasks: 100 Visibility: public Spec: Harbor task… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/repo2rlenv-code-instruct.n<1K0 likes513 downloads7h agoHugging Face07ed001 /ds-coder-instruct-v1 Dataset Card for DS Coder Instruct Dataset DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python. The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.imagetext-generation10K<n<100K5 likes369 downloads3y agoHugging Face08wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes255 downloads2y agoHugging Face09ed001 /ds-coder-instruct-v2 Dataset Card for DS Coder Instruct v2 Dataset Changes from v1: Added WizardLM evol data science samples Removed R samples from v2 DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2). The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.tabulartext-generation10K<n<100K13 likes230 downloads3y agoHugging Face10Dahoas /code-review-instruct-critique-revision Dataset Card for "code-review-instruct-critique-revision" More Information needed text10K<n<100K4 likes214 downloads4y agoHugging Face11CodeDevX /Vibe-Coding-Instructtexttext-generation1M<n<10M189 likes212 downloads3mo agoHugging Face12AtlasUnified /Code-Instruct-Setstext100K<n<1M6 likes205 downloads3y agoHugging Face13harman /deepcoder-train-deepcoder-qwen4b-instruct-cont-temp0_6-32k-hsrun_step230-codeonly_truncationtext10K<n<100K0 likes204 downloads11mo agoHugging Face14DCAgent2 /dcagent-dev-set-71-tasks-qwen-qwen3-coder-30b-a3b-instruct-20251117-231142textn<1K0 likes202 downloads10mo agoHugging Face15OALL /details_Qwen__Qwen2.5-Coder-14B-Instruct Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B-Instruct. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-14B-Instruct.tabular100K<n<1M0 likes188 downloads2y agoHugging Face16thesven /CodeMaster-Phi-Instruct Code Master Phi is a compiled dataset designed for training Phi3 instruct models. This dataset is focused on code-based data and integrates multiple high-quality sources to ensure a robust training foundation. The sources include: Replete-AI/code_bagel: A diverse collection of code snippets and examples. nickrosh/Evol-Instruct-Code-80k-v1: A dataset featuring evolved instructions for code generation tasks. iamtarun/python_code_instructions_18k_alpaca: A compilation of Python code… See the full description on the dataset page: https://huggingface.co/datasets/thesven/CodeMaster-Phi-Instruct.texttext-generation1M<n<10M0 likes180 downloads2y agoHugging Face17MaLA-LM /code-instruct-finalThis is a curated collection of code instruction tuning datasets that have been formatted in the LLAMA chat format and using markdown for code snippets. It subsets for the languages we seek to continue pretraining the MaLA-LM models on (refer to MaLA-LM/stack-final) using guesslang. The instruction tuning datasets we draw from are: ise-uiuc/Magicoder-OSS-Instruct-75K ise-uiuc/Magicoder-Evol-Instruct-110K glaiveai/glaive-code-assistant-v3 nuprl/EditPackFT-Multi likaixin/InstructCoder… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/code-instruct-final.text1M<n<10M3 likes174 downloads2y agoHugging Face18nyu-dice-lab /allenai_WildChat-1M-Full-neuralmagic_DeepSeek-Coder-V2-Instruct-FP8gatedtext100K<n<1M0 likes159 downloads2y agoHugging Face19open-llm-leaderboard-old /details_ajibawa-2023__Code-290k-6.7B-Instruct Dataset Card for Evaluation run of ajibawa-2023/Code-290k-6.7B-Instruct Dataset automatically created during the evaluation run of model ajibawa-2023/Code-290k-6.7B-Instruct on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__Code-290k-6.7B-Instruct.0 likes152 downloads3y agoHugging Face20open-llm-leaderboard-old /details_deepseek-ai__deepseek-coder-1.3b-instruct Dataset Card for Evaluation run of deepseek-ai/deepseek-coder-1.3b-instruct Dataset Summary Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-coder-1.3b-instruct on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_deepseek-ai__deepseek-coder-1.3b-instruct.0 likes145 downloads3y agoHugging Face21open-llm-leaderboard-old /details_TheBloke__CodeLlama-13B-Instruct-fp16 Dataset Card for Evaluation run of TheBloke/CodeLlama-13B-Instruct-fp16 Dataset Summary Dataset automatically created during the evaluation run of model TheBloke/CodeLlama-13B-Instruct-fp16 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__CodeLlama-13B-Instruct-fp16.0 likes135 downloads3y agoHugging Face22mlfoundations-dev /distill_r1_code_evol_instructtext1K<n<10K0 likes134 downloads2y agoHugging Face23CodeDevX /Vibe-Coding-Instruct-V2 Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/CodeDevX/Vibe-Coding-Instruct-V2.texttext-classification1M<n<10M11 likes133 downloads3mo agoHugging Face24open-llm-leaderboard-old /details_GeorgiaTechResearchInstitute__starcoder-gpteacher-code-instruct Dataset Card for Evaluation run of GeorgiaTechResearchInstitute/starcoder-gpteacher-code-instruct Dataset Summary Dataset automatically created during the evaluation run of model GeorgiaTechResearchInstitute/starcoder-gpteacher-code-instruct on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_GeorgiaTechResearchInstitute__starcoder-gpteacher-code-instruct.2 likes132 downloads3y agoHugging Face25jiachenli-ucsb /self-oss-instruct-sc2-exec-filter-prompt-codes-test-50ktext10K<n<100K0 likes130 downloads2y agoHugging Face26DCAgent2 /dcagent-dev-set-71-tasks-qwen-qwen3-coder-30b-a3b-instruct-20251116-070538textn<1K0 likes124 downloads10mo agoHugging Face27liodon-ai /gemma4-code-review-instruct gemma4-code-review-instruct 197K code review examples — 58K with chain-of-thought <think> reasoning traces. Built to train models that don't just flag issues, but explain their reasoning before delivering a review. Drop-in ready for SFT with any chat model. Why This Dataset Most code review datasets give you diff → comment. This one gives you diff → think → comment for 30% of examples — reasoning traces that show how to analyze a diff before writing the review.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/gemma4-code-review-instruct.texttext-generation100K<n<1M4 likes123 downloads3mo agoHugging Face28vikp /evol_instruct_code_filtered_39k Dataset Card for "evol_instruct_code_filtered_38k" Filtered version of nickrosh/Evol-Instruct-Code-80k-v1, with manual filtering, and automatic filtering based on quality and learning value classifiers. tabular10K<n<100K3 likes119 downloads3y agoHugging Face29kunishou /amenokaku-code-instruct Amenokaku-Code-Instruct Update: 2023/12/27データセットに JaxTon , プロになるJava のコードデータ 180 レコードを追加しました。 概要 コードに特化した5.2KのInstructionデータセットです。 データセットに含まれるデータは商用利用できるラインセンスが付与されたプログラミング学習コンテンツから収集、加工し作成しました(英語のコンテンツは日本語に自動翻訳し、翻訳の不自然な箇所を手動で修正)。 また、ライセンスが明記されていない学習コンテンツについては権利者に個別に連絡を取り、本データセットへの掲載の許諾を得ております。 データセット詳細 指示タスクの内訳としてはコード生成(code_generation)が1050レコード、コードの挙動確認(check_code_behavor)が150レコード、コードのバグ修正(code_fix)が4000レコードになります。 詳細な内訳は以下の通りになります。 source name… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/amenokaku-code-instruct.text1K<n<10K17 likes114 downloads2y agoHugging Face30OALL /details_Qwen__Qwen2.5-Coder-7B-Instruct Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-7B-Instruct.tabular100K<n<1M0 likes110 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.