datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
textbook_quality_programming
Dataset Card for "textbook_quality_programming"
Synthetic programming textbooks generated with GPT-3.5 and retrieval. Very high quality, aimed at being used in a phi replication. Currently 115M tokens. Covers many languages and technologies, with a bias towards python.
~10k of the books (65M tokens) use an older generation method, and average 6k tokens in length. ~1.5k books (50M tokens) use a newer generation method, with a more detailed outline, and average 33k tokens in… See the full description on the dataset page: https://huggingface.co/datasets/vikp/textbook_quality_programming.Competitive-Programmingpaloma_programming_languagesIndustryCorpus2_computer_programming_code
IndustryCorpus2: Programming
This repository contains the IndustryCorpus2: Programming domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year = {2024}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_computer_programming_code.programming_books_llama
Dataset Card for "programming_books_llama"
400M tokens of programming books generated by gpt-3.5 (70M tokens) and a finetuned codellama 34b. The gpt-3.5 data is extremely high quality. The llama data has lower quality and shorter length, but is still good. This was generated with the textbook quality repo.
PersonaSignal-PerceivabilityTest-Programming-Expertise-gpt-5-mini
Dataset card for PersonaSignal-PerceivabilityTest-Programming-Expertise-gpt-5-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"dimension_name": "programming_expertise",
"dimension_values": [
"Novice",
"Intermediate",
"Advanced"
],
"dimension_description": "Represents the user's practical fluency in software engineering. It shapes how they decompose problems, choose abstractions, weigh… See the full description on the dataset page: https://huggingface.co/datasets/JasonYan777/PersonaSignal-PerceivabilityTest-Programming-Expertise-gpt-5-mini.PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini
Dataset card for PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"dimension_name": "programming_expertise",
"dimension_values": [
"Novice",
"Intermediate",
"Advanced"
],
"dimension_description": "Represents the user's practical fluency in software engineering. It shapes how they decompose problems, choose abstractions, weigh… See the full description on the dataset page: https://huggingface.co/datasets/JasonYan777/PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini.python-programming-instructionspython_programming_questionslinear-programmingCode-290k-labels-programming_languages-NO_Chatgpt
Para etiquetar los lenguajes de programación en un conjunto de datos extenso de fragmentos de código, se aplicaron técnicas automatizadas de procesamiento de texto y patrones específicos de cada lenguaje, sin recurrir al uso de modelos de lenguaje avanzados como ChatGPT o LLMs. Se inició con la extracción y preparación de datos usando pandas, una biblioteca de análisis de datos en Python, que facilitó la manipulación y el procesamiento del conjunto de datos obtenido de Hugging Face's… See the full description on the dataset page: https://huggingface.co/datasets/NickyNicky/Code-290k-labels-programming_languages-NO_Chatgpt.quantum-compilation-and-programming
Neura Parse — Quantum Compilation & Programming
A code-heavy vertical on the quantum software/compilation stack: turning abstract quantum circuits and unitaries into device-executable programs. Covers unitary decomposition and circuit synthesis (Euler/ZYZ, KAK/Cartan, Solovay-Kitaev, Ross-Selinger gridsynth, numerical synthesis with BQSKit), gate-set/basis transpilation to native gate sets, qubit layout/mapping and routing under connectivity constraints (SABRE, VF2, SWAP… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-compilation-and-programming.spatial-programmingColumns:
input_text — text prompt (may be null)
input_image — reference image (may be null); at least one of text / image is set
code — Python source that builds the asset
glb — the resulting GLB file, raw bytes
platform — e.g. blender
type — e.g. modeling
model — which model wrote the code (fable, fable 5.1, opus 5, astra, sol)
Rows are appended one parquet shard per upload under data/.
competitive-programming-curated-600
🚀 Competitive Programming & Algorithmic Reasoning (Verbose CoT Reasoning)
This dataset contains 600 curated training records with in-depth, verbose 4-phase <Thinking> Chain-of-Thought reasoning, 100 frozen evaluation benchmark samples, and 50 frozen regression verification samples formatted in standard ChatML (messages) and Prompt-Target pairs, strictly following the Pioneer / Prometheus research paper 3-slice curriculum design.
📊 Dataset Composition & 3-Slice… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/competitive-programming-curated-600.multiturn_programming_binarized
Dataset Card for "multiturn_programming_binarized"
More Information needed
arjoonn-codechef-competitive-programming-ChatGPT4ointensive-programmingstackexchange-programming-cs-ctx-4096PersonaSignal-LeakageCheck-Programming-Expertise-claude-sonnet-4-5-20250929PersonaSignal-LeakageCheck-Programming-Expertise-gpt-4ostackexchange-programming-cs-ctx-2048stackexchange-programming-cs-ctx-6144PersonaSignal-LeakageCheck-Programming-Expertise-DPO-TinkerPersonaSignal-LeakageCheck-Programming-Expertise-Meta-Llama-3.1-8B-Instruct-Turbostackexchange-programming-cs-ctx-8192programming-languages-keywords
Dataset Card for "programming-languages-keywords"
Structured version of https://github.com/e3b0c442/keywords
Generated using:
r = requests.get("https://raw.githubusercontent.com/e3b0c442/keywords/main/README.md")
keywords = r.text.split("### ")[1:]
keywords = [i for i in keywords if not i.startswith("Sources")]
keywords = {i.split("\n")[0]:[j for j in re.findall("[a-zA-Z]*", i.split("\n",1)[1]) if j] for i in keywords}
keywords =… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/programming-languages-keywords.PersonaSignal-LeakageCheck-Programming-Expertise-gpt-4o-minispotlight-vikp-textbook_quality_programming-enrichment
Dataset Card for "spotlight-vikp-textbook_quality_programming-enrichment"
More Information needed
python_programming_QAturkish-programming-languages
