datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🌌 Omni-Frontier Distillation SFT
The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise
Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
"The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/TypeSafeAI/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.icd10pcs-coding-mcq
ICD-10-PCS Coding MCQ
405 multiple-choice items on ICD-10-PCS inpatient procedure coding — whether a model can
build a seven-character procedure code from documentation it is handed: root operation
selection, the seven character axes, Index→Tables verification, approach, device and qualifier
values, and the Official Guidelines.
Labels are what GPT-5.6-sol ruled they are. Items were written by Claude and adjudicated
by GPT-5.6-sol; where the two disagreed, the adjudicator's… See the full description on the dataset page: https://huggingface.co/datasets/chongpangnasilemak/icd10pcs-coding-mcq.agentic_coding_dataset
Agentic Coding Dataset
This dataset is a compilation of various coding and instruction-following datasets, designed to train agentic coding models.
Sources
This dataset aggregates samples from the following sources:
CodeAlpaca-20k
Instruction-following coding tasks.
Evol-CodeAlpaca-v1
Complex evolved coding instructions (WizardCoder style).
Code Review Instruct
Python code review, critique, and revision examples.
APPS (Automated Programming Progress Standard)… See the full description on the dataset page: https://huggingface.co/datasets/ethanker/agentic_coding_dataset.icd10cm-coding-mcq
ICD-10-CM Coding MCQ
403 multiple-choice items on ICD-10-CM diagnosis coding — whether a model can apply the
classification's conventions to documentation it is handed: Excludes1 and Excludes2 notes,
7th-character selection, placeholder X, laterality, Index→Tabular verification, combination
codes, specificity.
Labels are what GPT-5.6-sol ruled they are. Items were written by Claude and adjudicated
by GPT-5.6-sol; where the two disagreed, the adjudicator's ruling settled the… See the full description on the dataset page: https://huggingface.co/datasets/chongpangnasilemak/icd10cm-coding-mcq.visual-cortex-coding-qaAgentic-Coding-Tessa
Agentic Coding Dataset for Tessa
A comprehensive dataset for training coding agents with tool-use, reasoning, and software engineering capabilities.
Dataset Composition
This dataset combines multiple high-quality sources:
hermes_reasoning (20.0%): Tool-use and reasoning dataset - interstellarninja/hermes_reasoning_tool_use
search_arena (15.0%): Search and retrieval tasks - lmarena-ai/search-arena-24k
arena_human_pref (15.0%): Human preference data for alignment -… See the full description on the dataset page: https://huggingface.co/datasets/smirki/Agentic-Coding-Tessa.coding-variant
GenBench CoCG QA Dataset
Multi-hop genetic reasoning QA items generated from GenBench's knowledge
graph (Ensembl, ClinVar, VEP, BioGRID, STRING, Reactome, UniProt, GO,
SIGNOR, OmniPath, KEGG, DisGeNET, OpenTargets, PubTator3, GTEx, and more),
built for CoCG (Co-Evolving Confidence Graph) agent training.
2513 items across 11 task types.
Task types
task_type
count
coding_variant
53
conservation_reasoning
246
counterfactual
246
disease_reasoning
246… See the full description on the dataset page: https://huggingface.co/datasets/genbench-iitp/coding-variant.genbench-coding-qa
GenBench CoCG QA Dataset
Multi-hop genetic reasoning QA items generated from GenBench's knowledge
graph (Ensembl, ClinVar, VEP, BioGRID, STRING, Reactome, UniProt, GO,
SIGNOR, OmniPath, KEGG, DisGeNET, OpenTargets, PubTator3, GTEx, and more),
built for CoCG (Co-Evolving Confidence Graph) agent training.
8159 items across 11 task types.
Task types
task_type
count
coding_variant
159
conservation_reasoning
800
counterfactual
800
disease_reasoning
800… See the full description on the dataset page: https://huggingface.co/datasets/genbench-iitp/genbench-coding-qa.coding-model-rendered-qa
Rendered QA Dataset: Code & Text (700K)
Instruction-tuning dataset with optional rendered images for vision-language models.
Sources
Source
Samples
Has Context Image
OpenCoder Stage 2
436K
educational_instruct only
InstructCoder
108K
Yes (code input)
OpenOrca
200K
No (text-only)
Schema
Column
Type
Description
prompt
string
Instruction/question
prompt_image
Image?
Rendered prompt (optional)
context
string?
Code context… See the full description on the dataset page: https://huggingface.co/datasets/mustavinsu/coding-model-rendered-qa.
