datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
students-coding-questions-from-ai-assistant
Dataset Documentation
Overview
This dataset contains 6776 questions asked by students from CodeAid, an AI coding assistant, during a C programming class over a 12-week semester from January to April 2023. The course did not allow the use of ChatGPT, but CodeAid was permitted. CodeAid, powered by GPT-3, did not directly disclose code solutions even when requested by students. Instead, it functioned like a teaching assistant, providing scaffolded responses in natural… See the full description on the dataset page: https://huggingface.co/datasets/majeedkazemi/students-coding-questions-from-ai-assistant.context-ucurve-coding-agents
Context U-curve: 36 coding-agent runs under six context-clearing policies
How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report
"Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents"
(Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668).
A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.tech-debt-ai-coding
Debt Behind the AI Boom — Replication Data
Data for the paper:
Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
Yue Liu, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, Junkai Chen, David Lo
📄 arXiv:2603.28592 · 💻 Code: github.com/yueyueL/tech-debt-ai-coding
We mined 302.6K AI-authored commits from 6,299 GitHub repositories across five
AI coding assistants (GitHub Copilot, Claude, Cursor, Gemini, Devin), ran static
analysis… See the full description on the dataset page: https://huggingface.co/datasets/yueyuel/tech-debt-ai-coding.Solana-blockchain-360-CodingThis dataset contains 360 coding and tech related samples for the Solana blockchain.
Language: English
Coding-languages: Rust, Typescript, & C#
214 general knowledge samples
146 coding knowledge samples
Dataset Catalog:
201 Solana blockchain knowledge samples
49 Solana typescript coding samples
86 Solana rust coding samples
13 Solnet SDK knowledge samples
11 Solana c# coding samples
AI-Coding-Models
Dataset Card for 2026 AI Coding Models
Last Updated: 24 May 2026
Curated By: Joy Larkin
Language(s) (NLP): English
License: MIT
Repository: https://github.com/joylarkin/AI-Coding-Landscape
Blog: https://cleverhack.com/ai-coding-landscape
Dataset Description
CSV file of AI Coding Models released in 2026 & 2025.
Bwenge
Dataset Description
This dataset was created to develop a machine translation model for bidirectional translation between Kinyarwanda and English for education-based sentences, in particular for the Atingi learning platform.
Repository:link to the GitHub repository containing the code for training the model on this data, and the code for the collection of the monolingual data.
Data Format: TSV
Model: huggingface model link.
Dataset Summary
Data… See the full description on the dataset page: https://huggingface.co/datasets/Coding-With-Bashir/Bwenge.AI-Coding-Tools
Dataset Card for 2026 AI Coding Tools
Last Updated: 24 May 2026
Curated By: Joy Larkin
Language(s) (NLP): English
License: MIT
Repository: https://github.com/joylarkin/AI-Coding-Landscape
Blog: https://cleverhack.com/ai-coding-landscape
Dataset Description
CSV file of AI Coding Tools released in 2026 & 2025.
claude_opus-4.6_4.7_coding_reasoningclinical-quad-signal-detection-drift-ae-coding-variance-unblinding-risk-dsmb-decision-delay-v0.1
Clinical Quad: Signal Detection Drift × AE Coding Variance × Unblinding Risk × DSMB Decision Delay
This dataset targets safety governance collapse.
Signals weaken or shift.AE coding diverges across sites.Unblinding pressure rises.The DSMB response slows.
The quad can turn a manageable safety issue into a governance failure.
Variables
signal_detection_drift (low | medium | high)
ae_coding_variance (low | medium | high)
unblinding_risk (low | medium | high)… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-signal-detection-drift-ae-coding-variance-unblinding-risk-dsmb-decision-delay-v0.1.Coding_Sequences_13_Speciescoding_harmless_prompts
Coding Harmless Prompts
Benign coding and technical prompts for the harmless side of infosec refusal-direction extraction.
Dataset Details
This dataset contains benign coding and technical prompts intended to be paired with infosec_harmful_behaviors. The contrast helps isolate malicious coding intent rather than a general coding or technical-domain direction.
Rows:
train: 400
test: 120
Schema:
text: prompt string
Intended Use
Use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/zaakirio/coding_harmless_prompts.coding_exercisesfinance_codingVibe-Coding-Dataset
Vibe Coding Dataset
This dataset was published on Kaggle by Samyakraj Bayar and mirrored here.
Download
The dataset is available as a ZIP archive: Vibe Coding Dataset.zip
License
MIT
codingameThis dataset is scraped from codingame. Please check them out.
It contains coding problems with description, example input and output as well as about 5 test cases and 5 validator test cases per problem.
Because of the nature of the test cases defining only input and output, this is completly language agnostic and automatically verifiable.
This dataset is under Creative Commons Attribution Share Alike 3.0 license.
The dataset contains different types of puzzles:
Game Type
Number of… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/codingame.anesthesia-coding-31.5kFaqsfaq1Clinical_Coding_Finetuning_Formattedtemp_real_estate_yang_coding_studycoding-assistance-preferences
Coding Assitance Preferences
Coding Assistance Preferences is a dataset designed to study how programmers evaluate human and AI-generated responses to Python questions.
Each example presents a StackOverflow question, alongside two answers:
one written by a human on the forum and upvoted by users,
and one generated by an LLM (Gemini-2.0-Flash).
Annotators rated which response they preferred, the type of question (theoretical or practical), whether the responses suggest the same… See the full description on the dataset page: https://huggingface.co/datasets/NaomiDerel/coding-assistance-preferences.codingArt-chatnirvash-coding-healingjaguar-raw-datacodingajayvisit-with-us-datasetcodingArt-chatV2coding_challengecrypto-coding
