datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
context-ucurve-coding-agents
Context U-curve: 36 coding-agent runs under six context-clearing policies
How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report
"Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents"
(Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668).
A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.tech-debt-ai-coding
Debt Behind the AI Boom — Replication Data
Data for the paper:
Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
Yue Liu, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, Junkai Chen, David Lo
📄 arXiv:2603.28592 · 💻 Code: github.com/yueyueL/tech-debt-ai-coding
We mined 302.6K AI-authored commits from 6,299 GitHub repositories across five
AI coding assistants (GitHub Copilot, Claude, Cursor, Gemini, Devin), ran static
analysis… See the full description on the dataset page: https://huggingface.co/datasets/yueyuel/tech-debt-ai-coding.Bwenge
Dataset Description
This dataset was created to develop a machine translation model for bidirectional translation between Kinyarwanda and English for education-based sentences, in particular for the Atingi learning platform.
Repository:link to the GitHub repository containing the code for training the model on this data, and the code for the collection of the monolingual data.
Data Format: TSV
Model: huggingface model link.
Dataset Summary
Data… See the full description on the dataset page: https://huggingface.co/datasets/Coding-With-Bashir/Bwenge.coding_exercisesVibe-Coding-Dataset
Vibe Coding Dataset
This dataset was published on Kaggle by Samyakraj Bayar and mirrored here.
Download
The dataset is available as a ZIP archive: Vibe Coding Dataset.zip
License
MIT
CoDIGIT-EdinburghThis is an experimental, creative and participatory effort to implement a co-designed, situated and small-scale fine-tuning dataset for Large Language Models (LLMs).
The dataset contains 1197 items:
• 60 gender-oriented, co-designed prompts for sentence completion;
• 180 model completions generated by LLaMA 3.1 8B (three for each prompt);
• 897 gender bias scores, assigned to model responses by participants based on the experimental gender bias scale reported below;
• 60… See the full description on the dataset page: https://huggingface.co/datasets/elledilara/CoDIGIT-Edinburgh.faq1temp_real_estate_yang_coding_studycoding-assistance-preferences
Coding Assitance Preferences
Coding Assistance Preferences is a dataset designed to study how programmers evaluate human and AI-generated responses to Python questions.
Each example presents a StackOverflow question, alongside two answers:
one written by a human on the forum and upvoted by users,
and one generated by an LLM (Gemini-2.0-Flash).
Annotators rated which response they preferred, the type of question (theoretical or practical), whether the responses suggest the same… See the full description on the dataset page: https://huggingface.co/datasets/NaomiDerel/coding-assistance-preferences.jaguar-raw-datavisit-with-us-dataset
