CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M3 likes3.9k downloads23d agoHugging Face02Manusagents /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🌌 Omni-Frontier Distillation SFT The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection "The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.texttext-generation10M<n<100M6 likes1.8k downloads2mo agoHugging Face03greghavens /gpt-5.6-sol-coding-and-debugging-traces GPT-5.6 Sol Coding & Debugging Traces Verified software-engineering, independent model-judging, seed-authoring, defensive-security, and training-harness trajectories from GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an autonomous coding agent. Sessions show the observable development loop: inspecting repositories, reproducing failures, explaining evidence, editing files, running compilers and test suites, correcting mistakes, and verifying the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/gpt-5.6-sol-coding-and-debugging-traces.text-generation10K<n<100K44 likes1.1k downloads2mo agoHugging Face04Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes513 downloads23d agoHugging Face05Nobody05 /gpt-5.6-sol-coding-and-debugging-traces GPT-5.6 Sol Coding & Debugging Traces Verified software-engineering, independent model-judging, seed-authoring, defensive-security, and training-harness trajectories from GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an autonomous coding agent. Sessions show the observable development loop: inspecting repositories, reproducing failures, explaining evidence, editing files, running compilers and test suites, correcting mistakes, and verifying the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/gpt-5.6-sol-coding-and-debugging-traces.text-generation10K<n<100K1 likes112 downloads2mo agoHugging Face06himanshunakrani9 /mimo-coding-synthetic-5k MiMo Coding Synthetic 5.4K MiMo Coding Synthetic 5.4K is a purely synthetic coding instruction dataset generated with Xiaomi MiMo mimo-v2.5-pro. It contains 5,411 validated examples across programming languages, coding task types, and difficulty levels. The dataset is provided in two formats: A canonical rich JSONL format with metadata and labels. An OpenAI chat messages JSONL format for supervised fine-tuning pipelines. The generation run used 20 parallel workers for roughly… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/mimo-coding-synthetic-5k.texttext-generation10K<n<100K0 likes96 downloads3mo agoHugging Face07chongpangnasilemak /icd10pcs-coding-mcq ICD-10-PCS Coding MCQ 405 multiple-choice items on ICD-10-PCS inpatient procedure coding — whether a model can build a seven-character procedure code from documentation it is handed: root operation selection, the seven character axes, Index→Tables verification, approach, device and qualifier values, and the Official Guidelines. Labels are what GPT-5.6-sol ruled they are. Items were written by Claude and adjudicated by GPT-5.6-sol; where the two disagreed, the adjudicator's… See the full description on the dataset page: https://huggingface.co/datasets/chongpangnasilemak/icd10pcs-coding-mcq.textquestion-answeringn<1K0 likes80 downloads27d agoHugging Face08Voidreaper2026 /coding-master-dataset Coding Master Dataset Overview A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated. Records: 766,987 Format: JSONL / ShareGPT License: Apache 2.0 Sources CodeX-2M-Thinking (430,542 records) python-code-dataset-500k (559,515 records) StackPulse high-quality subset (20,205 records) CodeFeedback-Filtered-Instruction (156,525 records) secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.texttext-generation100K<n<1M3 likes77 downloads2mo agoHugging Face09chongpangnasilemak /icd10cm-coding-mcq ICD-10-CM Coding MCQ 403 multiple-choice items on ICD-10-CM diagnosis coding — whether a model can apply the classification's conventions to documentation it is handed: Excludes1 and Excludes2 notes, 7th-character selection, placeholder X, laterality, Index→Tabular verification, combination codes, specificity. Labels are what GPT-5.6-sol ruled they are. Items were written by Claude and adjudicated by GPT-5.6-sol; where the two disagreed, the adjudicator's ruling settled the… See the full description on the dataset page: https://huggingface.co/datasets/chongpangnasilemak/icd10cm-coding-mcq.textquestion-answeringn<1K0 likes68 downloads27d agoHugging Face10Manusagents /gpt-5.6-sol-coding-and-debugging-traces GPT-5.6 Sol Coding & Debugging Traces Verified software-engineering, independent model-judging, seed-authoring, defensive-security, and training-harness trajectories from GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an autonomous coding agent. Sessions show the observable development loop: inspecting repositories, reproducing failures, explaining evidence, editing files, running compilers and test suites, correcting mistakes, and verifying the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/gpt-5.6-sol-coding-and-debugging-traces.text-generation10K<n<100K1 likes67 downloads2mo agoHugging Face11ethanker /agentic_coding_dataset Agentic Coding Dataset This dataset is a compilation of various coding and instruction-following datasets, designed to train agentic coding models. Sources This dataset aggregates samples from the following sources: CodeAlpaca-20k Instruction-following coding tasks. Evol-CodeAlpaca-v1 Complex evolved coding instructions (WizardCoder style). Code Review Instruct Python code review, critique, and revision examples. APPS (Automated Programming Progress Standard)… See the full description on the dataset page: https://huggingface.co/datasets/ethanker/agentic_coding_dataset.texttext-generation100K<n<1M7 likes66 downloads10mo agoHugging Face12genbench-iitp /coding-variant GenBench CoCG QA Dataset Multi-hop genetic reasoning QA items generated from GenBench's knowledge graph (Ensembl, ClinVar, VEP, BioGRID, STRING, Reactome, UniProt, GO, SIGNOR, OmniPath, KEGG, DisGeNET, OpenTargets, PubTator3, GTEx, and more), built for CoCG (Co-Evolving Confidence Graph) agent training. 2513 items across 11 task types. Task types task_type count coding_variant 53 conservation_reasoning 246 counterfactual 246 disease_reasoning 246… See the full description on the dataset page: https://huggingface.co/datasets/genbench-iitp/coding-variant.tabularquestion-answering1K<n<10K1 likes60 downloads1mo agoHugging Face13smirki /Agentic-Coding-Tessa Agentic Coding Dataset for Tessa A comprehensive dataset for training coding agents with tool-use, reasoning, and software engineering capabilities. Dataset Composition This dataset combines multiple high-quality sources: hermes_reasoning (20.0%): Tool-use and reasoning dataset - interstellarninja/hermes_reasoning_tool_use search_arena (15.0%): Search and retrieval tasks - lmarena-ai/search-arena-24k arena_human_pref (15.0%): Human preference data for alignment -… See the full description on the dataset page: https://huggingface.co/datasets/smirki/Agentic-Coding-Tessa.texttext-generation10K<n<100K13 likes49 downloads1y agoHugging Face14OpceanAI /sota-codingtexttext-generation100K<n<1M1 likes45 downloads4mo agoHugging Face15asnelt /visual-cortex-coding-qatextquestion-answering100K<n<1M0 likes43 downloads10mo agoHugging Face16convaiinnovations /bilingual-coding-qa-dataset 🌐 Bilingual Coding Q&A Dataset 📊 Dataset Description A comprehensive bilingual (English-Hindi) dataset containing 25,151 high-quality question-answer pairsfocused on programming concepts, particularly Python, machine learning, and AI. This dataset was used to fine-tune coding assistant models and contains over 7 million tokens of training data. Dataset Statistics Metric Value Total Examples 25,151 Q&A pairs Total Lines 250,320+… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/bilingual-coding-qa-dataset.question-answering10K<n<100K1 likes42 downloads11mo agoHugging Face17cogbuji /MrGrammaticalOntology_clinical_coding Mr. Grammatical Ontology: Clinical Coding This dataset was created from a motivation to train Medical Large Language Models for improved fluency in clinical coding, as measurable by MedConceptsQA, an open-source medical coding evaluation benchmark designed to evaluate the understanding and reasoning capabilities of LLMs on medical concepts. It was extracted from the Centers for Medicare & Medicaid Services' International Classification of Diseases, Tenth Revision, Clinical… See the full description on the dataset page: https://huggingface.co/datasets/cogbuji/MrGrammaticalOntology_clinical_coding.textquestion-answering100K<n<1M2 likes33 downloads2y agoHugging Face18genbench-iitp /genbench-coding-qagated GenBench CoCG QA Dataset Multi-hop genetic reasoning QA items generated from GenBench's knowledge graph (Ensembl, ClinVar, VEP, BioGRID, STRING, Reactome, UniProt, GO, SIGNOR, OmniPath, KEGG, DisGeNET, OpenTargets, PubTator3, GTEx, and more), built for CoCG (Co-Evolving Confidence Graph) agent training. 8159 items across 11 task types. Task types task_type count coding_variant 159 conservation_reasoning 800 counterfactual 800 disease_reasoning 800… See the full description on the dataset page: https://huggingface.co/datasets/genbench-iitp/genbench-coding-qa.tabularquestion-answering1K<n<10K1 likes21 downloads1mo agoHugging Face19GG13412 /CodingQuestionDatabaseCodeLlamaThe questions, responses, and topics were generated with the codellama 7b model. They may be empty data points due to the AI generation. Main.json: CodeLlama 7b moneywordmath.json: pplx-7b-chat (Math Word Problems) textquestion-answeringn<1K0 likes16 downloads2y agoHugging Face20thunder-research-group /SNU_Thunder-synthetic-codinggated Dataset Card for SNU Thunder Synthetic Coding Dataset Summary This dataset was used as part of the post-training corpus for SnuLLM(to_fill). This dataset consists of Korean and English question-answer pairs. Questions are sourced from publicly available datasets, and answers were generated using open large language models (Exaone 3.5, LLaMA 3.3, Qwen 2.5). It is intended for research and non-commercial use. Supported Tasks Tasks: Python coding Languages… See the full description on the dataset page: https://huggingface.co/datasets/thunder-research-group/SNU_Thunder-synthetic-coding.textquestion-answering1K<n<10K2 likes10 downloads1y agoHugging Face21mustavinsu /coding-model-rendered-qa Rendered QA Dataset: Code & Text (700K) Instruction-tuning dataset with optional rendered images for vision-language models. Sources Source Samples Has Context Image OpenCoder Stage 2 436K educational_instruct only InstructCoder 108K Yes (code input) OpenOrca 200K No (text-only) Schema Column Type Description prompt string Instruction/question prompt_image Image? Rendered prompt (optional) context string? Code context… See the full description on the dataset page: https://huggingface.co/datasets/mustavinsu/coding-model-rendered-qa.imagequestion-answering100K<n<1M0 likes8 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.