datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
devopsbench-100
DevOpsBench-100
DevOpsBench-100 is a synthetic long-horizon software-engineering / SRE agent
benchmark: 100 tasks over one executable world ("NovaCart", a mid-size
e-commerce SaaS) with 72 SQLite tables,
1451 seeded rows, a 38-file monorepo with 417 commits,
and 97 MCP tools spanning a first-party engineering stack
(tickets, PRs, CI, deployments, canaries, migrations, feature flags, metrics,
alerts, incidents, chat, knowledge base) plus deliberately disagreeing
vendor-shaped… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/devopsbench-100.stack-v3-devops
The Stack v3 DevOps Corpus
13,234,862 complete infrastructure units extracted from
The Stack v3,
grouped into seven classes and gated on content rather than popularity.
A unit is not a file, it is the thing an engineer would actually run: a Helm chart
arrives with its Chart.yaml, values.yaml and every template; a Terraform module
with all of its .tf files; an Ansible role with its tasks, defaults and handlers.
That is only possible because The Stack v3 groups rows by repository… See the full description on the dataset page: https://huggingface.co/datasets/Helmcode/stack-v3-devops.omnimcp_devops_cloud_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_devops_cloud_teaser.devops-cloud-instruction-dataset
DevOps & Cloud Infrastructure Dataset
Professional instruction-response pairs for DevOps engineers covering Kubernetes, Docker, Terraform, CI/CD, and cloud services (AWS, Azure).
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Devops Cloud topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/devops-cloud-instruction-dataset.devops-kubernetes-sft-100k
DevOps and Kubernetes SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality DevOps and Kubernetes conversations designed to train AI assistants capable of supporting platform engineers, SREs, and DevOps practitioners.
Dataset Description
This dataset covers production-grade Kubernetes operations, cloud infrastructure, CI/CD pipelines, GitOps workflows, and platform engineering across 13 specialized categories. Each record follows the ShareGPT… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/devops-kubernetes-sft-100k.devops-sft-dataset
DevOps SFT Instruction Dataset
This dataset contains 8,076 high-quality instruction-response pairs specifically generated for fine-tuning a DevOps domain-specialized language model. It was used in the Supervised Fine-Tuning (SFT) phase of the Ulysses model training pipeline.
Dataset Description
Instructions were generated using the Gemini API (gemini-2.0-flash) and Ollama (qwen2.5-coder:7b) by feeding chunks of official DevOps documentation and GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/jalpan04/devops-sft-dataset.devops-kubernetes-iac-sft-dpo-2026
⚙️ Enterprise DevOps AI, Kubernetes SRE & IaC SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step SRE root-cause Chain-of-Thought (<thought>) diagnostic trees for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Senior Site Reliability Engineers (SRE), Principal Cloud Architects, and DevSecOps Specialists.
📊 Dataset Architecture & Highlights
Multi-Turn SRE… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/devops-kubernetes-iac-sft-dpo-2026.bio-devops-synthetic-instructions
Bio-DevOps Synthetic Instructions
This dataset contains synthetic instruction-following examples for biomedical-style data-engineering and scientific-computing workflows.
It was created for educational and portfolio use as part of a LoRA/QLoRA fine-tuning project using Qwen/Qwen2.5-Coder-7B-Instruct.
Related model:
AiLLMBS/qwen25-coder-bio-devops-lora
Dataset Contents
The dataset includes synthetic examples for:
Python CSV validation
pandas duplicate checks
bash… See the full description on the dataset page: https://huggingface.co/datasets/AiLLMBS/bio-devops-synthetic-instructions.devops-promql-sre-curated-600
🚀 DevOps SRE PromQL Telemetry Diagnostics & Alert Triage
This dataset contains 600 curated training records with in-depth, verbose 4-phase <Thinking> Chain-of-Thought reasoning, 100 frozen evaluation benchmark samples, and 50 frozen regression verification samples formatted in standard ChatML (messages) and Prompt-Target pairs, strictly following the Pioneer / Prometheus research paper 3-slice curriculum design.
📊 Dataset Composition & 3-Slice Breakdown… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/devops-promql-sre-curated-600.devops-v1
Dataset Card for Docker & Kubernetes Troubleshooting Dataset
Dataset Summary
This dataset contains 256 comprehensive question-solution pairs covering Docker and Kubernetes troubleshooting scenarios. It includes common issues, error messages, and their detailed solutions for container orchestration and cloud-native infrastructure management. The dataset spans Docker fundamentals, Kubernetes core concepts, cloud-managed Kubernetes services (EKS, AKS), and advanced… See the full description on the dataset page: https://huggingface.co/datasets/pavanmantha/devops-v1.mirror-devops-sft-dataset
DevOps SFT Instruction Dataset
This dataset contains 8,076 high-quality instruction-response pairs specifically generated for fine-tuning a DevOps domain-specialized language model. It was used in the Supervised Fine-Tuning (SFT) phase of the Ulysses model training pipeline.
Dataset Description
Instructions were generated using the Gemini API (gemini-2.0-flash) and Ollama (qwen2.5-coder:7b) by feeding chunks of official DevOps documentation and GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-devops-sft-dataset.smolified-backend-and-devops-buddy
🤏 smolified-backend-and-devops-buddy
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model draganite/smolified-backend-and-devops-buddy.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 68bd3135)
Records: 1890
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by draganite.
Generated via Smolify.ai.
