CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anupambayen /AnupamB-Coder-Dataset AnupamB-Coder-Dataset A large-scale synthetic dataset of Python and SQL examples spanning basic to expert difficulty — purpose-built for training AnupamB-Coder-110M, a GPT-style code language model built entirely from scratch on a gaming laptop. The Story Behind This Dataset Most code datasets on HuggingFace come from scraping GitHub or StackOverflow. This one is different. Every single example in this dataset was generated by a pure Python template engine — no GPT, no… See the full description on the dataset page: https://huggingface.co/datasets/anupambayen/AnupamB-Coder-Dataset.texttext-generation10M<n<100M1 likes430 downloads6mo agoHugging Face02ed001 /ds-coder-instruct-v1 Dataset Card for DS Coder Instruct Dataset DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python. The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.imagetext-generation10K<n<100K5 likes369 downloads3y agoHugging Face03Lite-Coder /LiteCoder-Terminal-SFT LiteCoder-SFT-Terminal Paper | Code | Blog Post LiteCoder-SFT-Terminal is a dataset of 11,255 agent trajectories in terminal environments, introduced in the paper LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents. Fine-tuned on this data, the LiteCoder-Terminal-30b-a3b-sft model achieves 31.5% Pass@1 on Terminal Bench Pro, while the LiteCoder-Terminal-4b-sft model shows distinct gains over its baseline. Released Artifacts… See the full description on the dataset page: https://huggingface.co/datasets/Lite-Coder/LiteCoder-Terminal-SFT.texttext-generation10K<n<100K9 likes333 downloads4mo agoHugging Face04ed001 /ds-coder-instruct-v2 Dataset Card for DS Coder Instruct v2 Dataset Changes from v1: Added WizardLM evol data science samples Removed R samples from v2 DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2). The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.tabulartext-generation10K<n<100K13 likes230 downloads3y agoHugging Face05JohnBeanerson /strudel-coder Claude Code session traces for JohnBeanerson/strudel-coder This dataset contains redacted Claude Code session traces collected while working on https://github.com/ultralazr/strudel-coder.git. The traces were exported with cc-share-hf and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each file in the repo root is a redacted Claude Code session in its native JSONL format (one entry per line). HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/JohnBeanerson/strudel-coder.tabulartext-generationn<1K0 likes164 downloads4mo agoHugging Face06smcleod /golang-coderQ&A style combined, deduplicated dataset including portions of: Golang best practices and coding guides (general Q&A) https://huggingface.co/datasets/smcleod/golang-programming-style-best-practices (MIT) Golang questions (general Q&A) https://huggingface.co/datasets/ExAi/Code-Golang-QA-2k (Apache2) Golang functions (code & description) https://huggingface.co/datasets/google/code_x_glue_ct_code_to_text (c-uda) Golang snippets (code & description)… See the full description on the dataset page: https://huggingface.co/datasets/smcleod/golang-coder.texttext-generation100K<n<1M19 likes139 downloads2y agoHugging Face07code-rag-bench /humanevalHumanEval dataset annotated with the ground-truth programming solutions, to enable evaluations for retrieval and retrieval augmented code generation. Please refer to code-rag-becnch for more details. texttext-generationn<1K0 likes103 downloads2y agoHugging Face08AmareshHebbar /cpt-coder-sft CPT / HCPCS Procedure Coder Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Procedure descriptions → correct CPT/HCPCS code with RVU data Why download this Build procedure coding assistants, verify CPT code assignments, or automate outpatient charge capture. Covers all specialties in the CMS PFS. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/cpt-coder-sft.texttext-generation10K<n<100K0 likes100 downloads3mo agoHugging Face09thetemirbolatov /TILO.RA_CODER_Dataset TILO.RA CODER Dataset Объединённый русско-английский датасет для обучения и поиска по коду. Формат — пары question / code: вопрос на естественном языке → готовый код-ответ. Датасет собран для локального ассистента TILO.RA CODER — офлайн-помощника по программированию Скачать по ссылке https://github.com/thetemirbolatov/TILO.RA_CODER_Dataset/releases/download/v1.0.0/tilora_knowledge_merged.jsonl Состав Источник Язык Записей English coding… See the full description on the dataset page: https://huggingface.co/datasets/thetemirbolatov/TILO.RA_CODER_Dataset.texttext-generation100K<n<1M1 likes99 downloads8d agoHugging Face10WithinUsAI /Python_GOD_Coder_Omniforge_AI_12k Python GOD Coder Omniforge AI 12k Creator: Within Us AI A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist. This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model: implementation with tests strict code-only instruction following debugging and repair refactoring for readability and production readiness next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.texttext-generation10K<n<100K1 likes94 downloads7mo agoHugging Face11AmareshHebbar /icd10-coder-sft ICD-10-CM Medical Coder Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Maps clinical descriptions to ICD-10-CM codes Why download this Fine-tune LLMs to automatically assign ICD-10-CM codes from clinical text. Useful for EHR automation, medical coding assistants, and clinical NLP pipelines. Dataset stats… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/icd10-coder-sft.texttext-generation10K<n<100K0 likes94 downloads3mo agoHugging Face12VatsaDev /code-reviewA Scrape of the codereview stack exchange, good for high quality code texttext-generation10K<n<100K3 likes91 downloads3y agoHugging Face13Nellyw888 /RTL-Coder_7b_reasoning_tb_combined Verireason-RTL-Coder_7b_reasoning_tb_combined For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation This is the combined version of VeriReason-RTL-Coder_7b_reasoning_tb and VeriReason-RTL-Coder_7b_reasoning_tb_simple. Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_combined Project… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning_tb_combined.texttext-generation1K<n<10K0 likes83 downloads1y agoHugging Face14bhavin-gajjar /laravel-coder-lora-train Laravel Coder LoRA — end-to-end training data Instruction-tuning datasets for Laravel Bob (Qwen / CodeLlama / DeepSeek modelfiles).Built from official Laravel docs v10.x–v13.x plus version-detection examples. Each config has a single schema — do not mix raw (Alpaca) with chat / lora_* (messages). Configs Config Path Rows Schema raw raw/laravel_training.jsonl 1253 instruction, input, output, topic, version chat chat/laravel_training_chat.jsonl 1253… See the full description on the dataset page: https://huggingface.co/datasets/bhavin-gajjar/laravel-coder-lora-train.texttext-generation1K<n<10K1 likes83 downloads3mo agoHugging Face15prithivMLmods /Coder-Stat Coder-Stat Dataset Overview The Coder-Stat dataset is a collection of programming-related data, including problem IDs, programming languages, original statuses, and source code snippets. This dataset is designed to assist in the analysis of coding patterns, error types, and performance metrics. Dataset Details Modalities Tabular: The dataset is structured in a tabular format. Text: Contains text data, including source code snippets. Formats… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Coder-Stat.tabulartext-classification10K<n<100K3 likes76 downloads2y agoHugging Face16Nellyw888 /VeriReason-RTL-Coder_7b_reasoning_tb Verireason-RTL-Coder_7b_reasoning_tb For implementation details, visit our GitHub repository: VeriReason and our page Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb.texttext-generation1K<n<10K4 likes74 downloads1y agoHugging Face17code-rag-bench /mbppMBPP dataset annotated with ground-truth programming solutions, to enable evaluations for retrieval and retrieval-augmented code generation. Please refer to code-rag-bench for more details. texttext-generationn<1K1 likes73 downloads2y agoHugging Face18CoderBak /minif2f minif2f Dataset The minif2f dataset is a collection of mathematical problems and their formal statements, designed for formal mathematics and theorem proving tasks. Dataset Description Dataset Summary The minif2f dataset contains mathematical problems from various sources (like AMC competitions) along with their formal statements in the Lean theorem prover format. Each example includes both informal mathematical statements and their corresponding formal… See the full description on the dataset page: https://huggingface.co/datasets/CoderBak/minif2f.texttext-generationn<1K0 likes59 downloads10mo agoHugging Face19Nellyw888 /RTL-Coder_7b_reasoning Verireason-RTL-Coder_7b_reasoning_tb_simple For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning.texttext-generation1K<n<10K1 likes58 downloads1y agoHugging Face20code-rag-bench /odexODEX dataset annotated with the ground-truth library documentation, to enable evaluations for retrieval and retrieval-augmented code generation. Please refer to [code-rag-bench] for more details. texttext-generationn<1K0 likes57 downloads2y agoHugging Face21Nellyw888 /VeriReason-RTL-Coder_7b_reasoning_tb_simple Verireason-RTL-Coder_7b_reasoning_tb_simple For implementation details, visit our GitHub repository: VeriReason and our page Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.texttext-generationn<1K0 likes57 downloads1y agoHugging Face22xlelords /orbis-coder Orbis Coder Dataset (10K) A coding-first instruction dataset to train or fine-tune assistants that behave like Orbis Coder — friendly, concise, practical, and focused on helping people build, debug, and ship software. This dataset is intended for: instruction-tuning / SFT LoRA / QLoRA fine-tunes “persona + skill” alignment for coding assistants quick experiments + dataset viewer testing What this dataset contains Most rows are coding help across many languages and… See the full description on the dataset page: https://huggingface.co/datasets/xlelords/orbis-coder.texttext-generation10K<n<100K0 likes52 downloads8mo agoHugging Face23matorus /coderTraining dataset for finetuning for human-eval. This dataset has been created from the following datasets: sahil2801/CodeAlpaca-20k sahil2801/code_instructions_120k mhhmm/leetcode-solutions-python teknium1/GPTeacher Script for generating dataset: create_dataset.py. texttext-generation10K<n<100K5 likes50 downloads3y agoHugging Face24Nellyw888 /RTL-Coder_small RTL-Coder_small For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to enhance the performance of pre-trained… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_small.texttext-generation1K<n<10K1 likes49 downloads1y agoHugging Face25Jackrong /qwen3-coder-480b-distill-mini qwen3-coder-480b-distill-mini Short Description This dataset is distilled using Qwen3-Coder-480B-A35B-Instruct.We extracted 10,000 code questions from microsoft/rStar-Coder as seed problems, distilled them with 32K context, and after cleaning and filtering, 9,543 samples remain.License: Apache-2.0. Dataset Overview Seed Source: 10,000 code reasoning problems sampled from microsoft/rStar-Coder. Distillation Model: Qwen3-Coder-480B-A35B-Instruct (480B… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/qwen3-coder-480b-distill-mini.texttext-classification1K<n<10K14 likes46 downloads1y agoHugging Face26delimitter /synoema-coder-3b-tools-corpus Synoema Tools — Training Corpora Exact corpora used to fine-tune the 100% Synoema agentic tool-use models (3B, 1.5B). Website: https://synoema.tech Files File Used for Examples merged_seq_c8.jsonl 3B C8 (100%) 18317 merged_seq_c12.jsonl 1.5B C12 (100%) 17321 targeted/targeted_seq_c9mw_3b.jsonl 3B multi-write fix (TU4/TU13) 44 targeted/targeted_seq_c11fix_1.5b.jsonl 1.5B fix (TU4/TU13/TU20/TU30) 36 targeted/targeted_seq_c10fix_0.8b.jsonl 0.8B fix… See the full description on the dataset page: https://huggingface.co/datasets/delimitter/synoema-coder-3b-tools-corpus.texttext-generation10K<n<100K0 likes41 downloads4mo agoHugging Face27domofon /Gasai-Agent-CoderForge Gasai-Agent-CoderForge 5,205 verified-resolved agentic software-engineering trajectories, re-authored into the Gasai harness format for pretraining small language models. Derived from togethercomputer/CoderForge-Preview (reward==1, Python). Each row is one complete tool-use trace serialized as a single Gasai control-token sequence in the gasai field. Sister dataset to GasaiAI/Gasai-Agent-5k (from nvidia/Open-SWE-Traces) — byte-identical format, same pipeline. What… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Gasai-Agent-CoderForge.texttext-generation1K<n<10K0 likes38 downloads3mo agoHugging Face28hornsan1 /Vera-Agentic-Coder Vera Agentic Coder - GLM-5.2 REAP-50 Calibration Corpus Public-safe calibration corpus for a steep 50% REAP expert-pruning pass on GLM-5.2. The mix is weighted to preserve agent harness/tool behavior and coding capability, while retaining baseline general knowledge, basic math/reasoning, infra, security, shell, and home-lab competence. This is intended for router/expert activation profiling only. It is not a supervised training set or benchmark. Stats… See the full description on the dataset page: https://huggingface.co/datasets/hornsan1/Vera-Agentic-Coder.texttext-generation10K<n<100K0 likes36 downloads3mo agoHugging Face29code-rag-bench /ds1000DS-1000 dataset annotated with the ground-truth library documentation, to enable evaluations for retrieval and retrieval-augmented code generation. Please refer to [code-rag-bench] for more details texttext-generation1K<n<10K1 likes35 downloads2y agoHugging Face30AmareshHebbar /radiology-coder-sft Radiology Report ICD-10 Coder Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Radiology reports / impressions → ICD-10-CM codes for all documented findings Why download this Automate radiology coding, build report-to-code pipelines for RIS/PACS integration, or train models to extract diagnosis codes from chest X-ray… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/radiology-coder-sft.texttext-generation10K<n<100K0 likes30 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.