datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SciCode-Programming-Problems
DATA3: Programming Problems Generation Dataset
Dataset Overview
DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational methods in… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Programming-Problems.Vulnerable_Programming_DatasetVulnerable Programming Dataset
Overview
The Vulnerable Programming Dataset is a comprehensive collection of 550 unique code vulnerabilities across 10 programming languages: Python, JavaScript, PHP, Java, Ruby, Go, TypeScript, C++, SQL, and C. Designed for cybersecurity professionals, red teamers, pentesters, and developers, this dataset highlights unconventional vulnerabilities such as insecure interprocess communication, misconfigured rate limiting, insecure dependency pinning, and logic… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Vulnerable_Programming_Dataset.SciCode-Programming-Problems
DATA3: Programming Problems Generation Dataset
Dataset Overview
DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Programming-Problems.Haumea-ARC1-Programmatic-Reasoning
Haumea ARC-1 Training Solver
A modular, law-based solver for the Abstraction and Reasoning Corpus (ARC-1) dataset. This solver successfully solves 400/400 training tasks from the ARC-1 dataset using a systematic composition of geometric, topological, and logical "laws."
Overview
The solver is structured around a "Mega Engine" that applies a library of modular laws to solve complex visual reasoning tasks.
Solve Rate: 400/400 (ARC-1 Training Set)
Methodology: Modular Law… See the full description on the dataset page: https://huggingface.co/datasets/oncloudai/Haumea-ARC1-Programmatic-Reasoning.imabari_wiki_qa_v4_program_validated
Imabari QA v4 — Program Validated
Dataset Summary
Imabari QA v4 — Program Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
It is derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The dataset contains synthetic question-answer pairs together with generated reasoning traces in the thinking field.
The primary difference from the original Imabari Wiki QA v4 dataset is the… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_program_validated.quantum-compilation-and-programming
Neura Parse — Quantum Compilation & Programming
A code-heavy vertical on the quantum software/compilation stack: turning abstract quantum circuits and unitaries into device-executable programs. Covers unitary decomposition and circuit synthesis (Euler/ZYZ, KAK/Cartan, Solovay-Kitaev, Ross-Selinger gridsynth, numerical synthesis with BQSKit), gate-set/basis transpilation to native gate sets, qubit layout/mapping and routing under connectivity constraints (SABRE, VF2, SWAP… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-compilation-and-programming.DoD-Instruction-6055-17-Emergency-Management-Program
🚨 DoD Emergency Management Program
Maintainer: Terry Eppler
Owner: US Federal Government
Source: DoD Instruction 6055.17
Source Version: Change 4, effective December 1, 2025
Ownership of Source: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Emergency Management Program Question-Answer Dataset is a structured, document-grounded natural-language dataset derived from DoD Instruction 6055.17, “DoD Emergency… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-6055-17-Emergency-Management-Program.Vulnerable_Programming_DatasetVulnerable Programming Dataset
Overview
The Vulnerable Programming Dataset is a comprehensive collection of 550 unique code vulnerabilities across 10 programming languages: Python, JavaScript, PHP, Java, Ruby, Go, TypeScript, C++, SQL, and C. Designed for cybersecurity professionals, red teamers, pentesters, and developers, this dataset highlights unconventional vulnerabilities such as insecure interprocess communication, misconfigured rate limiting, insecure dependency pinning, and logic… See the full description on the dataset page: https://huggingface.co/datasets/bharath-nuthalapati/Vulnerable_Programming_Dataset.DOD-Instruction-1000-25-Personnel-Identity-Protection-Program
DoD Personnel Identity Protection Program
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on Department of Defense Instruction 1000.25, “DoD Personnel Identity Protection Program,” dated March 2, 2016.
The instruction establishes policy, governance, and organizational responsibilities for the Department of Defense Personnel Identity Protection Program. It also… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Instruction-1000-25-Personnel-Identity-Protection-Program.DoD-Instruction-5200-01-Information-Security-Program
DoD Information Security and SCI Protection Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on DoD Instruction 5200.01, “DoD Information Security Program and Protection of Sensitive Compartmented Information (SCI),” dated April 21, 2016, and incorporating Change 2 effective October 1, 2020.
The source establishes the overarching Department… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-5200-01-Information-Security-Program.Program-Management-Improvement-Accountability-Act-of-2016
Dataset Description
Maintainer: Terry Eppler
Owner: US Federal Government
The **Program Management Improvement Accountability Act of 2016 ** is an English-language instructional dataset containing 150 substantive question-and-answer records derived from the Program Management Improvement Accountability Act of 2016.
The Act, commonly abbreviated as PMIAA, was enacted as Public Law 114-264 on December 14, 2016. It amended title 31 of the United States Code to strengthen… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Program-Management-Improvement-Accountability-Act-of-2016.Vulnerable_Programming_DatasetVulnerable Programming Dataset
Overview
The Vulnerable Programming Dataset is a comprehensive collection of 550 unique code vulnerabilities across 10 programming languages: Python, JavaScript, PHP, Java, Ruby, Go, TypeScript, C++, SQL, and C. Designed for cybersecurity professionals, red teamers, pentesters, and developers, this dataset highlights unconventional vulnerabilities such as insecure interprocess communication, misconfigured rate limiting, insecure dependency pinning, and logic… See the full description on the dataset page: https://huggingface.co/datasets/VIGNESHWARAN04/Vulnerable_Programming_Dataset.Synthetic_Java_Dialog_And_Programs_LLM_TrainingThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
Java Programming Examples Dataset
Dataset Description
This dataset contains 8 distinct Java programs with 10 conversational examples each, synthetically generated from a larger dataset of 80+ programs. Each program has 10,000 variants, providing a diverse set of Java code examples covering various programming… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_Java_Dialog_And_Programs_LLM_Training.mojo-programming-language-qnaA synthetic dataset from Claude Sonnet 3.5. The source documents are real pulled from the Mojo documentation, but everything else is synthetic.
ai-programming-cookbook
AI Programming Recipes Dataset
A Q&A dataset aimed at providing in-depth recipes on building complex AI systems for LLM fine-tuning. All the original data comes from open-source documentations.
Preprocessing
Multiple preprocessing steps were used to make the dataset ready for LoRa fine-tuning:
Convert the ipynb/rst files to markdown as it will be our output format. This was done thanks to nbconvert/pydantic.
Use an LLM (DeepSeek V3.1) to unclutter the output cells of… See the full description on the dataset page: https://huggingface.co/datasets/paulprt/ai-programming-cookbook.aurora_programmer_data
My Awesome Dataset
A comprehensive description of my awesome dataset.
Dataset Description
This dataset contains images of cats and dogs. The images were collected from [mention data source(s), e.g., a specific website, scraped from the internet]. It is intended for use in image classification tasks. The dataset consists of [number] images, with approximately [percentage]% allocated to the training set and [percentage]% to the test set. [Add more details about the… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/aurora_programmer_data.beir_cqadupstack_programmers_test
beir_cqadupstack_programmers_test
BEIR CQADupStack/programmers test split
Field
Value
Benchmark
beir
Sub-benchmark
cqadupstack_programmers
Type
retrieval
Items
876
Exported from Langfuse.
usda-program-comparison
USDA Program Comparison
72 USDA vs FHA vs Conventional comparison scenarios.
Details
Records: 72
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Tate Thompson, NMLS #2473962
Publisher: Thompson Mortgage Group
Thompson Alpha Logic
Side-by-side USDA vs FHA vs Conventional analysis across 72 buyer scenarios. Calculates monthly payment, total cost, and 5-year breakeven for each program by credit score, income, and location — with… See the full description on the dataset page: https://huggingface.co/datasets/wendymthompson/usda-program-comparison.dpa-programs-2026
DPA Programs 2026
34 down payment assistance programs across 13 states.
Details
Records: 34
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Beau Thompson, NMLS #1615561
Publisher: Good News Lending
Thompson Alpha Logic
State-by-state mapping of 2026 DPA grants cross-referenced with FHA/USDA eligibility to find 'Net-Zero Cash' purchase zones where buyers can close with $0 out of pocket by stacking DPA with zero-down loan programs.
This… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/dpa-programs-2026.dpa-programs-2026
Down Payment Assistance Programs 2026
Comprehensive database of down payment assistance (DPA) programs across 13 Southeast US states, curated by Good News Lending.
Dataset Description
34 active DPA programs from state Housing Finance Agencies (HFAs), including grants, forgivable loans, deferred loans, and repayable second mortgages.
States Covered (13)
Tennessee, Mississippi, Alabama, Georgia, Florida, Kentucky, Louisiana, North Carolina, South Carolina… See the full description on the dataset page: https://huggingface.co/datasets/wendymthompson/dpa-programs-2026.perl-programming-qaSt.Clair_Programsqa_program_modules_docs
