CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01microsoft /rStar-Coder rStar-Coder Dataset Project GitHub | Paper Dataset Description rStar-Coder is a large-scale competitive code problem dataset containing 418K programming problems, 580K long-reasoning solutions, and rich test cases of varying difficulty levels. This dataset aims to enhance code reasoning capabilities in large language models, particularly in handling competitive code problems. Experiments on Qwen models (1.5B-14B) across various code reasoning benchmarks demonstrate… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/rStar-Coder.text1M<n<10M246 likes53k downloads1y agoHugging Face02IFM /Code-Reasoning Code-Reasoning Dataset Description Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Code-Reasoning.texttext-generation100M<n<1B68 likes42k downloads20d agoHugging Face03CoderOfCode /ship-tracking-data2 likes22k downloads2mo agoHugging Face04togethercomputer /CoderForge-Preview CoderForge-Preview: SOTA Open Dataset for Training Efficient Agents CoderForge-Preview is the largest open test-verified coding agent dataset. Fine-tuning Qwen-3 32B on it, we boost SWE-Bench Verified performance 23.0% → 59.4% pass@1 and rank #1 among open-data and #2 among open-weight models ≤32B parameters. Limitations Adaptability to different scaffolds: We generated all trajectories using a single scaffold and fixed tool set (no permutations). Models trained via… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/CoderForge-Preview.text100K<n<1M176 likes5.6k downloads7mo agoHugging Face05code-rag-bench /github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos) text100K<n<1M2 likes3.3k downloads2y agoHugging Face06QuixiAI /dolphin-coder dolphin-coder This dataset is transformed from https://www.kaggle.com/datasets/erichartford/leetcode-rosetta it is used to train dolphin-coder model text100K<n<1M62 likes1.6k downloads3y agoHugging Face07IIGroup /X-Coder-SFT-376k X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests Dataset Overview X-Coder-SFT-376k is a large-scale, fully synthetic dataset for advancing competitive programming. The dataset comprises 4 subsets with a total of 887,321 synthetic records across 423,883 unique queries. It is designed for supervised fine-tuning and suitbale for cold start to train code reasoning foundations. X-Coder-SFT-376k is curated by sota reasoning models.… See the full description on the dataset page: https://huggingface.co/datasets/IIGroup/X-Coder-SFT-376k.texttext-generation100K<n<1M21 likes1.5k downloads8mo agoHugging Face08Lite-Coder /LiteCoder-Terminal-RL-preview LiteCoder-Terminal-RL-preview Paper | Code | Blog Post This dataset contains 602 standardized Harbor terminal environments and was released as part of the paper LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents. Unlike static text-only instructions, these environments are fully executable and are designed to support the training of terminal-based agents. Environment Generation Pipeline The lack of high-quality, executable… See the full description on the dataset page: https://huggingface.co/datasets/Lite-Coder/LiteCoder-Terminal-RL-preview.tabulartext-generationn<1K6 likes1.4k downloads3mo agoHugging Face09inclusionAI /Ling-Coder-SFT 🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub Ling-Coder Dataset The Ling-Coder Dataset comprises the following components: Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples. Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples. Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SFT.texttext-generation1M<n<10M45 likes1.4k downloads1y agoHugging Face10ronantakizawa /github-codereview Code Review Dataset A large-scale dataset of the best human-written code reviews from top GitHub repositories. Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response. The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable. This provides a natural signal for training models to: Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.tabulartext-generation100K<n<1M62 likes1.4k downloads7mo agoHugging Face11tsinghua-sigs-robot-lab /veriloop-coder-e1-evaluation-evidence VeriLoop Coder-E1 Evaluation Evidence This repository contains the public evaluation-evidence packages referenced by the official VeriLoop Coder-E1 benchmark result files. Model repository: tsinghua-sigs-robot-lab/veriloop-coder-e1 Evidence packages Benchmark Evidence directory DeepSWE veriloop-coder-e1-deepswe-evaluation-evidence-v1.0.0 SWE-bench Pro veriloop-coder-e1-swe-bench-pro-evaluation-evidence-v1.0.0 SWE-bench Verified… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/veriloop-coder-e1-evaluation-evidence.0 likes1.1k downloads2mo agoHugging Face12nebula2025 /CodeR-Pile Towards A Generalist Code Embedding Model Based On Massive Data Synthesis Introduction This repository contains the synthetic training data introduced in the paper Towards A Generalist Code Embedding Model Based On Massive Data Synthesis. The dataset is designed to enhance text embeddings for code retrieval tasks. For more details, please refer to our Github repo: CodeR. Load Dataset Simple Example An example to load the dataset:… See the full description on the dataset page: https://huggingface.co/datasets/nebula2025/CodeR-Pile.text1M<n<10M4 likes998 downloads11mo agoHugging Face13code-rag-bench /github-reposThe entire dump of GitHub repositories. text100K<n<1M2 likes825 downloads2y agoHugging Face14ricdomolm /mini-coder-trajs-400kGenerated using Qwen 3 Coder 30B A3B, mini-swe-agent, and SWE-smith. Used to train the mini-coder models Citation @article{olmedo2026computational, title={Computational Arbitrage in AI Model Markets}, author={Olmedo, Ricardo and Sch{\"o}lkopf, Bernhard and Hardt, Moritz}, journal={The International Conference on Machine Learning}, year={2026} } text100K<n<1M16 likes814 downloads1mo agoHugging Face15AzerChakir /CodeReviewWithSummaryQAgatedtextn<1K0 likes696 downloads1mo agoHugging Face16Fraser /dream-coder Program Synthesis Data Generated program synthesis datasets used to train dreamcoder. Currently just supports text & list data. text1K<n<10K6 likes684 downloads4y agoHugging Face17inclusionAI /Ling-Coder-SyntheticQA 🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub Ling-Coder Dataset The Ling-Coder Dataset comprises the following components: Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples. Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples. Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SyntheticQA.texttext-generation10M<n<100M17 likes643 downloads1y agoHugging Face18gudo7208 /CAD-Coder CAD-Coder Dataset This is the official dataset for the paper "CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward". Accepted at NeurIPS 2025 (Poster) Dataset Description CAD-Coder Dataset is a large-scale Text-to-CadQuery dataset containing natural language descriptions of 3D CAD models paired with executable CadQuery Python code. The dataset enables training and evaluating language models to generate parametric CAD code from textual descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/gudo7208/CAD-Coder.texttext-generation100K<n<1M5 likes533 downloads9mo agoHugging Face19fan-shu /swe-mt-combined-coderforge-hero-lego-nex-swezero fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded in order and concatenated into a single config so one training epoch visits every trajectory exactly once (no interleave / no oversampling). Built from fan-shu/swe-instruct-trajectories-empty-think-inserted. Source subsets (7) togethercomputer__CoderForge-Preview nvidia__SWE-Zero-openhands-trajectories nex-agi__agent-sft… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero.text100K<n<1M0 likes514 downloads2mo agoHugging Face20Coder-AN /StreakNet-Dataset StreakNet-Dataset StreakNet-Dataset is an underwater laser imaging dataset for UCLR systems, introduced in the paper StreakNet-Arch: An Anti-scattering Network-based Architecture for Underwater Carrier LiDAR-Radar Imaging. It comprises a collection of streak-tube images captured by a UCLR system at distances of 10m, 13m, 15m, and 20m, contributing 2,695,168 real-world underwater 3D point cloud data. For the associated source code, models, and comprehensive usage instructions… See the full description on the dataset page: https://huggingface.co/datasets/Coder-AN/StreakNet-Dataset.image-to-3d0 likes475 downloads1y agoHugging Face21coderchen01 /MMSD2.0 MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System This is a copy of the dataset uploaded on Hugging Face for easy access. The original data comes from this work, which is an improvement upon a previous study. Usage from typing import TypedDict, cast import pytorch_lightning as pl from datasets import Dataset, load_dataset from torch import Tensor from torch.utils.data import DataLoader from transformers import CLIPProcessor class… See the full description on the dataset page: https://huggingface.co/datasets/coderchen01/MMSD2.0.imagefeature-extraction10K<n<100K7 likes454 downloads2y agoHugging Face22coderofpears /clanker-data0 likes454 downloads2mo agoHugging Face23anupambayen /AnupamB-Coder-Dataset AnupamB-Coder-Dataset A large-scale synthetic dataset of Python and SQL examples spanning basic to expert difficulty — purpose-built for training AnupamB-Coder-110M, a GPT-style code language model built entirely from scratch on a gaming laptop. The Story Behind This Dataset Most code datasets on HuggingFace come from scraping GitHub or StackOverflow. This one is different. Every single example in this dataset was generated by a pure Python template engine — no GPT, no… See the full description on the dataset page: https://huggingface.co/datasets/anupambayen/AnupamB-Coder-Dataset.texttext-generation10M<n<100M1 likes431 downloads6mo agoHugging Face24isthatshan /WestGenesis-Coder-SFT-100M Dataset Overview WestGenesis-Coder-Dataset is a meticulously curated coding dataset designed specifically for instruction-based model tuning and fine-tuning of existing models with enhanced code generation capabilities. This represents one of the largest and most comprehensively filtered corpora of publicly available coding data on the Hugging Face platform, with a non-thinking approach that emphasizes direct, concise code outputs for rapid model training. Key… See the full description on the dataset page: https://huggingface.co/datasets/isthatshan/WestGenesis-Coder-SFT-100M.texttext-generation1M<n<10M0 likes416 downloads3mo agoHugging Face25togethercomputer /CoderForge-Preview-32B-SWE-Bench-Verified-Evaluation-trajectoriestabularn<1K13 likes395 downloads8mo agoHugging Face26aysinghal /code-retrieval-training-datasettext100K<n<1M1 likes388 downloads6mo agoHugging Face27open-llm-leaderboard-old /details_Ramikan-BR__tinyllama_PY-CODER-4bit-lora_4k-v12 Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Ramikan-BR__tinyllama_PY-CODER-4bit-lora_4k-v12.text-generation10K<n<100K0 likes366 downloads2y agoHugging Face28ed001 /ds-coder-instruct-v1 Dataset Card for DS Coder Instruct Dataset DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python. The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.imagetext-generation10K<n<100K5 likes364 downloads3y agoHugging Face29IIGroup /X-Coder-RL-40k X-Coder-RL-40k X-Coder-RL-40k is a fully synthetic reinforcement learning dataset for competitive programming, containing 40k high-quality tasks with verified test cases. Dataset Structure The dataset is organized by difficulty level: File Difficulty part_0000.parquet Easiest part_0001.parquet Easy part_0002.parquet Medium part_0003.parquet Hard part_0004.parquet Hardest Task Difficulty Distribution Table: Distribution of Proprietary… See the full description on the dataset page: https://huggingface.co/datasets/IIGroup/X-Coder-RL-40k.textn<1K2 likes361 downloads8mo agoHugging Face30code-rag-bench /stackoverflow-postsThe StackOverflow posts retrieval source for code-rag-bench. text1M<n<10M2 likes351 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.