CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.4k downloads9mo agoHugging Face02codeparrot /self-instruct-starcoder Self-instruct-starcoder Summary Self-instruct-starcoder is a dataset that was generated by prompting starcoder to generate new instructions based on some human-written seed instructions. The underlying process is explained in the paper self-instruct. This algorithm gave birth to famous machine generated datasets such as Alpaca and Code Alpaca which are two datasets obtained by prompting OpenAI text-davinci-003 engine. Our approach While our method is… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/self-instruct-starcoder.text1K<n<10K64 likes690 downloads3y agoHugging Face03wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes255 downloads2y agoHugging Face04Dahoas /code-review-instruct-critique-revision Dataset Card for "code-review-instruct-critique-revision" More Information needed text10K<n<100K4 likes214 downloads4y agoHugging Face05harman /deepcoder-train-deepcoder-qwen4b-instruct-cont-temp0_6-32k-hsrun_step230-codeonly_truncationtext10K<n<100K0 likes204 downloads11mo agoHugging Face06DCAgent2 /dcagent-dev-set-71-tasks-qwen-qwen3-coder-30b-a3b-instruct-20251117-231142textn<1K0 likes202 downloads10mo agoHugging Face07OALL /details_Qwen__Qwen2.5-Coder-14B-Instruct Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B-Instruct. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-14B-Instruct.tabular100K<n<1M0 likes188 downloads2y agoHugging Face08thesven /CodeMaster-Phi-Instruct Code Master Phi is a compiled dataset designed for training Phi3 instruct models. This dataset is focused on code-based data and integrates multiple high-quality sources to ensure a robust training foundation. The sources include: Replete-AI/code_bagel: A diverse collection of code snippets and examples. nickrosh/Evol-Instruct-Code-80k-v1: A dataset featuring evolved instructions for code generation tasks. iamtarun/python_code_instructions_18k_alpaca: A compilation of Python code… See the full description on the dataset page: https://huggingface.co/datasets/thesven/CodeMaster-Phi-Instruct.texttext-generation1M<n<10M0 likes180 downloads2y agoHugging Face09MaLA-LM /code-instruct-finalThis is a curated collection of code instruction tuning datasets that have been formatted in the LLAMA chat format and using markdown for code snippets. It subsets for the languages we seek to continue pretraining the MaLA-LM models on (refer to MaLA-LM/stack-final) using guesslang. The instruction tuning datasets we draw from are: ise-uiuc/Magicoder-OSS-Instruct-75K ise-uiuc/Magicoder-Evol-Instruct-110K glaiveai/glaive-code-assistant-v3 nuprl/EditPackFT-Multi likaixin/InstructCoder… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/code-instruct-final.text1M<n<10M3 likes174 downloads2y agoHugging Face10nyu-dice-lab /allenai_WildChat-1M-Full-neuralmagic_DeepSeek-Coder-V2-Instruct-FP8gatedtext100K<n<1M0 likes159 downloads2y agoHugging Face11mlfoundations-dev /distill_r1_code_evol_instructtext1K<n<10K0 likes134 downloads2y agoHugging Face12jiachenli-ucsb /self-oss-instruct-sc2-exec-filter-prompt-codes-test-50ktext10K<n<100K0 likes130 downloads2y agoHugging Face13DCAgent2 /dcagent-dev-set-71-tasks-qwen-qwen3-coder-30b-a3b-instruct-20251116-070538textn<1K0 likes124 downloads10mo agoHugging Face14liodon-ai /gemma4-code-review-instruct gemma4-code-review-instruct 197K code review examples — 58K with chain-of-thought <think> reasoning traces. Built to train models that don't just flag issues, but explain their reasoning before delivering a review. Drop-in ready for SFT with any chat model. Why This Dataset Most code review datasets give you diff → comment. This one gives you diff → think → comment for 30% of examples — reasoning traces that show how to analyze a diff before writing the review.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/gemma4-code-review-instruct.texttext-generation100K<n<1M4 likes123 downloads3mo agoHugging Face15vikp /evol_instruct_code_filtered_39k Dataset Card for "evol_instruct_code_filtered_38k" Filtered version of nickrosh/Evol-Instruct-Code-80k-v1, with manual filtering, and automatic filtering based on quality and learning value classifiers. tabular10K<n<100K3 likes119 downloads3y agoHugging Face16OALL /details_Qwen__Qwen2.5-Coder-7B-Instruct Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-7B-Instruct.tabular100K<n<1M0 likes110 downloads2y agoHugging Face17DCAgent2 /dcagent-dev-set-71-tasks-qwen-qwen3-coder-30b-a3b-instruct-20251111-105527textn<1K0 likes99 downloads10mo agoHugging Face18hubistrauss /tiny-codes-instructBase dataset: iamtarun/code_instructions_120k_alpaca text1M<n<10M0 likes97 downloads2y agoHugging Face19Dahoas /code-review-instruct-critique-revision-pythontext1K<n<10K10 likes88 downloads4y agoHugging Face20jtatman /python-github-code-instruct-filtered-5k Dataset Card for "python-github-code-instruct-filtered-5k" This fine dataset tomekkorbak/python-github-code, filtered by scores greater than 0.03. Feedback and additional columns generated through OpenAI and Cohere responses. texttext-generation1K<n<10K7 likes87 downloads2y agoHugging Face21DCAgent2 /dcagent-dev-set-71-tasks-qwen-qwen3-coder-30b-a3b-instruct-20251116-150649textn<1K0 likes77 downloads10mo agoHugging Face22DCAgent2 /swebench_verified_random_100_folders_Qwen3_Coder_480B_A35B_Instruct_FP8_20260429_192404textn<1K0 likes77 downloads5mo agoHugging Face23JulianAT /SynthUI-Code-Instruct-2k-v1Synth UI 🎹 https://www.synthui.design Dataset details This dataset aims to provide a diverse collection of NextJS code snippets, along with their corresponding instructions, to facilitate the training of language models for NextJS-related tasks. It is designed to cover a wide range of NextJS functionalities, including UI components, routing, state management, and more. This dataset consists of: Note: The dataset is seperated into two main parts: raw Contains only the… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/SynthUI-Code-Instruct-2k-v1.texttext-generation1K<n<10K0 likes76 downloads2y agoHugging Face24DCAgent2 /dcagent-dev-set-71-tasks-qwen-qwen3-coder-30b-a3b-instruct-20251115-190408textn<1K0 likes69 downloads10mo agoHugging Face25DONG19 /instruct_code_search_net Dataset Card for "instruct_code_search_net" More Information needed text1M<n<10M1 likes67 downloads3y agoHugging Face26stojchet /deepseek.coder.1.3b.instruct.python.mbpp.emptytextn<1K0 likes65 downloads2y agoHugging Face27DCAgent2 /dcagent-dev-set-71-tasks-qwen-qwen3-coder-30b-a3b-instruct-20251116-190724textn<1K0 likes64 downloads10mo agoHugging Face28erythropygia /Instruct-Python-Code-Turkish Dataset Card for Instruct-Python-Code-Turkish Language: Turkish Dataset Description The translation was performed using the Google translation model to ensure high-quality, accurate translation. Dataset Details Size: ≈5K Translation tool: Google Translate Data format: Instruct, Output texttext-generation1K<n<10K1 likes61 downloads2y agoHugging Face29mlfoundations-dev /a1_code_star_coder_instruct_eval_636d mlfoundations-dev/a1_code_star_coder_instruct_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 15.0 51.0 72.2 28.2 34.2 35.9 29.0 6.9 5.3 AIME24 Average Accuracy: 15.00% ± 0.85% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 13.33% 4 30 2 13.33% 4 30 3 16.67% 5 30 4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_star_coder_instruct_eval_636d.tabular1K<n<10K1 likes55 downloads1y agoHugging Face30DCAgent /eval-Qwen3-Coder-30B-A3B-Instruct_16concurrency_openhands_eval_c_terminal-bench-2.0textn<1K0 likes55 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.