datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stack-v3-devops
The Stack v3 DevOps Corpus
13,234,862 complete infrastructure units extracted from
The Stack v3,
grouped into seven classes and gated on content rather than popularity.
A unit is not a file, it is the thing an engineer would actually run: a Helm chart
arrives with its Chart.yaml, values.yaml and every template; a Terraform module
with all of its .tf files; an Ansible role with its tasks, defaults and handlers.
That is only possible because The Stack v3 groups rows by repository… See the full description on the dataset page: https://huggingface.co/datasets/Helmcode/stack-v3-devops.devops-kubernetes-iac-sft-dpo-2026
⚙️ Enterprise DevOps AI, Kubernetes SRE & IaC SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step SRE root-cause Chain-of-Thought (<thought>) diagnostic trees for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Senior Site Reliability Engineers (SRE), Principal Cloud Architects, and DevSecOps Specialists.
📊 Dataset Architecture & Highlights
Multi-Turn SRE… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/devops-kubernetes-iac-sft-dpo-2026.devops-predictive-logs
🔮 DevOps Predictive Logs Dataset
A synthetic dataset of realistic DevOps log sequences for training and benchmarking predictive failure models.
📊 Dataset Summary
This dataset contains realistic DevOps log scenarios covering common infrastructure failure patterns. Each log entry includes metadata about the failure scenario, severity, and time-to-failure, making it ideal for training predictive models.
Total Logs: ~150+ entriesScenarios: 10 unique failure patternsFormat:… See the full description on the dataset page: https://huggingface.co/datasets/Snaseem2026/devops-predictive-logs.deepfabric-devops-reasoning-traces
Dataset Description
DevOps reasoning traces
Dataset Details
Created by: Always Further
License: CC BY 4.0
Language(s): [English
Dataset Size: 10050
Data Splits
[train]
Dataset Creation
This dataset was created using DeepFabric, an open-source tool for generating high-quality training datasets for AI models.
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/nolabs/deepfabric-devops-reasoning-traces.stackexchange_devopssmoltrace-devops-tasks
SMOLTRACE Synthetic Dataset
This dataset was generated using the TraceMind MCP Server's synthetic data generation tools.
Dataset Info
Tasks: 100
Format: SMOLTRACE evaluation format
Generated: AI-powered synthetic task generation
Usage with SMOLTRACE
from datasets import load_dataset
# Load dataset
dataset = load_dataset("MCP-1st-Birthday/smoltrace-devops-tasks")
# Use with SMOLTRACE
# smoltrace-eval --model openai/gpt-4 --dataset-name… See the full description on the dataset page: https://huggingface.co/datasets/MCP-1st-Birthday/smoltrace-devops-tasks.devops-promql-sre-curated-600
🚀 DevOps SRE PromQL Telemetry Diagnostics & Alert Triage
This dataset contains 600 curated training records with in-depth, verbose 4-phase <Thinking> Chain-of-Thought reasoning, 100 frozen evaluation benchmark samples, and 50 frozen regression verification samples formatted in standard ChatML (messages) and Prompt-Target pairs, strictly following the Pioneer / Prometheus research paper 3-slice curriculum design.
📊 Dataset Composition & 3-Slice Breakdown… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/devops-promql-sre-curated-600.devops-v1
Dataset Card for Docker & Kubernetes Troubleshooting Dataset
Dataset Summary
This dataset contains 256 comprehensive question-solution pairs covering Docker and Kubernetes troubleshooting scenarios. It includes common issues, error messages, and their detailed solutions for container orchestration and cloud-native infrastructure management. The dataset spans Docker fundamentals, Kubernetes core concepts, cloud-managed Kubernetes services (EKS, AKS), and advanced… See the full description on the dataset page: https://huggingface.co/datasets/pavanmantha/devops-v1.devops_augdevops_sjmy-issues-dataset
Dataset Card for Dataset Name
Dataset Summary in English
This customized dataset is made of a corpus of commun Github issues, typically utilized for tracking bugs or features within a repositories. This self-constructed corpus can serve multiple purposes, such as analyzing the time taken to resolve open issues or pull requests, training a classifier to tag issues based on their descriptions (e.g., "bug," "enhancement," "question"), or developing a semantic search engine… See the full description on the dataset page: https://huggingface.co/datasets/devopsmarc/my-issues-dataset.PL-DevOps-Instructexp_8_8_domain_shift_devops_test25exp_8_8_domain_shift_devops_test5devops-training
Dataset Description
DevOps reasoning traces
Dataset Details
Created by: Always Further
License: CC BY 4.0
Language(s): [English
Dataset Size: 10050
Data Splits
[train]
Dataset Creation
This dataset was created using DeepFabric, an open-source tool for generating high-quality training datasets for AI models.
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/shahiddanial3545/devops-training.devops-guide-demodeepfabric-devops-with-toolsdevops-event-log-5000mlops-devops-sentiment
MLOps & DevOps Sentiment Dataset
Dataset description
A domain-specific sentiment dataset containing real-world MLOps and DevOps
scenarios labeled as POSITIVE or NEGATIVE. Built to fine-tune sentiment
classifiers for technical operations contexts where general-purpose models
(trained on movie reviews) underperform.
Why this dataset exists
General sentiment models misclassify technical sentences. For example:
"The pipeline failed silently" →… See the full description on the dataset page: https://huggingface.co/datasets/atulkrs/mlops-devops-sentiment.devops-kubectl-v1devops-v2DevOps_Gymsmolified-backend-and-devops-buddy
🤏 smolified-backend-and-devops-buddy
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model draganite/smolified-backend-and-devops-buddy.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 68bd3135)
Records: 1890
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by draganite.
Generated via Smolify.ai.
devops-lead-training-datadevops-event-log-augmentedDevOps_Gym_Oracle
