datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🌌 Omni-Frontier Distillation SFT
The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise
Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
"The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.gpt-5.6-sol-coding-and-debugging-traces
GPT-5.6 Sol Coding & Debugging Traces
Verified software-engineering, independent model-judging, seed-authoring,
defensive-security, and training-harness trajectories from
GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an
autonomous coding agent. Sessions show the observable development loop:
inspecting repositories, reproducing failures, explaining evidence, editing
files, running compilers and test suites, correcting mistakes, and verifying
the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/gpt-5.6-sol-coding-and-debugging-traces.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.gpt-5.6-sol-coding-and-debugging-traces
GPT-5.6 Sol Coding & Debugging Traces
Verified software-engineering, independent model-judging, seed-authoring,
defensive-security, and training-harness trajectories from
GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an
autonomous coding agent. Sessions show the observable development loop:
inspecting repositories, reproducing failures, explaining evidence, editing
files, running compilers and test suites, correcting mistakes, and verifying
the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/gpt-5.6-sol-coding-and-debugging-traces.mimo-coding-synthetic-5k
MiMo Coding Synthetic 5.4K
MiMo Coding Synthetic 5.4K is a purely synthetic coding instruction dataset generated with Xiaomi MiMo mimo-v2.5-pro.
It contains 5,411 validated examples across programming languages, coding task types, and difficulty levels. The dataset is provided in two formats:
A canonical rich JSONL format with metadata and labels.
An OpenAI chat messages JSONL format for supervised fine-tuning pipelines.
The generation run used 20 parallel workers for roughly… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/mimo-coding-synthetic-5k.icd10pcs-coding-mcq
ICD-10-PCS Coding MCQ
405 multiple-choice items on ICD-10-PCS inpatient procedure coding — whether a model can
build a seven-character procedure code from documentation it is handed: root operation
selection, the seven character axes, Index→Tables verification, approach, device and qualifier
values, and the Official Guidelines.
Labels are what GPT-5.6-sol ruled they are. Items were written by Claude and adjudicated
by GPT-5.6-sol; where the two disagreed, the adjudicator's… See the full description on the dataset page: https://huggingface.co/datasets/chongpangnasilemak/icd10pcs-coding-mcq.coding-master-dataset
Coding Master Dataset
Overview
A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated.
Records: 766,987
Format: JSONL / ShareGPT
License: Apache 2.0
Sources
CodeX-2M-Thinking (430,542 records)
python-code-dataset-500k (559,515 records)
StackPulse high-quality subset (20,205 records)
CodeFeedback-Filtered-Instruction (156,525 records)
secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.icd10cm-coding-mcq
ICD-10-CM Coding MCQ
403 multiple-choice items on ICD-10-CM diagnosis coding — whether a model can apply the
classification's conventions to documentation it is handed: Excludes1 and Excludes2 notes,
7th-character selection, placeholder X, laterality, Index→Tabular verification, combination
codes, specificity.
Labels are what GPT-5.6-sol ruled they are. Items were written by Claude and adjudicated
by GPT-5.6-sol; where the two disagreed, the adjudicator's ruling settled the… See the full description on the dataset page: https://huggingface.co/datasets/chongpangnasilemak/icd10cm-coding-mcq.gpt-5.6-sol-coding-and-debugging-traces
GPT-5.6 Sol Coding & Debugging Traces
Verified software-engineering, independent model-judging, seed-authoring,
defensive-security, and training-harness trajectories from
GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an
autonomous coding agent. Sessions show the observable development loop:
inspecting repositories, reproducing failures, explaining evidence, editing
files, running compilers and test suites, correcting mistakes, and verifying
the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/gpt-5.6-sol-coding-and-debugging-traces.agentic_coding_dataset
Agentic Coding Dataset
This dataset is a compilation of various coding and instruction-following datasets, designed to train agentic coding models.
Sources
This dataset aggregates samples from the following sources:
CodeAlpaca-20k
Instruction-following coding tasks.
Evol-CodeAlpaca-v1
Complex evolved coding instructions (WizardCoder style).
Code Review Instruct
Python code review, critique, and revision examples.
APPS (Automated Programming Progress Standard)… See the full description on the dataset page: https://huggingface.co/datasets/ethanker/agentic_coding_dataset.coding-variant
GenBench CoCG QA Dataset
Multi-hop genetic reasoning QA items generated from GenBench's knowledge
graph (Ensembl, ClinVar, VEP, BioGRID, STRING, Reactome, UniProt, GO,
SIGNOR, OmniPath, KEGG, DisGeNET, OpenTargets, PubTator3, GTEx, and more),
built for CoCG (Co-Evolving Confidence Graph) agent training.
2513 items across 11 task types.
Task types
task_type
count
coding_variant
53
conservation_reasoning
246
counterfactual
246
disease_reasoning
246… See the full description on the dataset page: https://huggingface.co/datasets/genbench-iitp/coding-variant.Agentic-Coding-Tessa
Agentic Coding Dataset for Tessa
A comprehensive dataset for training coding agents with tool-use, reasoning, and software engineering capabilities.
Dataset Composition
This dataset combines multiple high-quality sources:
hermes_reasoning (20.0%): Tool-use and reasoning dataset - interstellarninja/hermes_reasoning_tool_use
search_arena (15.0%): Search and retrieval tasks - lmarena-ai/search-arena-24k
arena_human_pref (15.0%): Human preference data for alignment -… See the full description on the dataset page: https://huggingface.co/datasets/smirki/Agentic-Coding-Tessa.sota-codingvisual-cortex-coding-qabilingual-coding-qa-dataset
🌐 Bilingual Coding Q&A Dataset
📊 Dataset Description
A comprehensive bilingual (English-Hindi) dataset containing 25,151 high-quality question-answer pairsfocused on programming concepts, particularly Python, machine learning, and AI. This dataset was used to fine-tune coding assistant models and contains over 7 million tokens of training data.
Dataset Statistics
Metric
Value
Total Examples
25,151 Q&A pairs
Total Lines
250,320+… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/bilingual-coding-qa-dataset.MrGrammaticalOntology_clinical_coding
Mr. Grammatical Ontology: Clinical Coding
This dataset was created from a motivation to train Medical Large Language Models for improved fluency in clinical
coding, as measurable by MedConceptsQA, an open-source medical coding
evaluation benchmark designed to evaluate the understanding and reasoning capabilities of LLMs on medical concepts. It was extracted from the Centers for Medicare & Medicaid Services'
International Classification of Diseases, Tenth Revision, Clinical… See the full description on the dataset page: https://huggingface.co/datasets/cogbuji/MrGrammaticalOntology_clinical_coding.genbench-coding-qa
GenBench CoCG QA Dataset
Multi-hop genetic reasoning QA items generated from GenBench's knowledge
graph (Ensembl, ClinVar, VEP, BioGRID, STRING, Reactome, UniProt, GO,
SIGNOR, OmniPath, KEGG, DisGeNET, OpenTargets, PubTator3, GTEx, and more),
built for CoCG (Co-Evolving Confidence Graph) agent training.
8159 items across 11 task types.
Task types
task_type
count
coding_variant
159
conservation_reasoning
800
counterfactual
800
disease_reasoning
800… See the full description on the dataset page: https://huggingface.co/datasets/genbench-iitp/genbench-coding-qa.CodingQuestionDatabaseCodeLlamaThe questions, responses, and topics were generated with the codellama 7b model.
They may be empty data points due to the AI generation.
Main.json: CodeLlama 7b
moneywordmath.json: pplx-7b-chat (Math Word Problems)
SNU_Thunder-synthetic-coding
Dataset Card for SNU Thunder Synthetic Coding
Dataset Summary
This dataset was used as part of the post-training corpus for SnuLLM(to_fill).
This dataset consists of Korean and English question-answer pairs. Questions are sourced from publicly available datasets, and answers were generated using open large language models (Exaone 3.5, LLaMA 3.3, Qwen 2.5). It is intended for research and non-commercial use.
Supported Tasks
Tasks: Python coding
Languages… See the full description on the dataset page: https://huggingface.co/datasets/thunder-research-group/SNU_Thunder-synthetic-coding.coding-model-rendered-qa
Rendered QA Dataset: Code & Text (700K)
Instruction-tuning dataset with optional rendered images for vision-language models.
Sources
Source
Samples
Has Context Image
OpenCoder Stage 2
436K
educational_instruct only
InstructCoder
108K
Yes (code input)
OpenOrca
200K
No (text-only)
Schema
Column
Type
Description
prompt
string
Instruction/question
prompt_image
Image?
Rendered prompt (optional)
context
string?
Code context… See the full description on the dataset page: https://huggingface.co/datasets/mustavinsu/coding-model-rendered-qa.
