datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ia-census
Internet Archive census (April 2016) — per-file md5+sha1
The April 2016 Internet Archive census
recorded md5 and sha1 for every file of every public IA item —
the only bulk file-hash inventory of the Archive ever published (the
census was not repeated; nothing uploaded after 2016-04-11 appears
here). This dataset is those two dumps joined — they are line-aligned
outputs of one item walk — into single witness rows co-observing both
digests, republished as page-indexed parquet for… See the full description on the dataset page: https://huggingface.co/datasets/david-ar/ia-census.ambig-iac
Ambig-IaC: Ambiguous Infrastructure-as-Code Benchmark
A benchmark dataset of 300 tasks for testing AI agents that generate Infrastructure-as-Code (Terraform) configurations from ambiguous natural language intents.
Project page: https://zyang37.github.io/ambig-iac.github.io/
Dataset Description
This dataset is sourced from IaC-Eval. We performed manual fixes to the original Terraform configurations and validated that all 300 tasks pass terraform plan. Each task also… See the full description on the dataset page: https://huggingface.co/datasets/znyang/ambig-iac.devops-kubernetes-iac-sft-dpo-2026
⚙️ Enterprise DevOps AI, Kubernetes SRE & IaC SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step SRE root-cause Chain-of-Thought (<thought>) diagnostic trees for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Senior Site Reliability Engineers (SRE), Principal Cloud Architects, and DevSecOps Specialists.
📊 Dataset Architecture & Highlights
Multi-Turn SRE… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/devops-kubernetes-iac-sft-dpo-2026.ambig-iac-sample
Ambig-IaC Random 50
A random subset of 50 rows sampled without replacement from all 300 rows of
the default/train split of znyang/ambig-iac.
Source revision: 96429693e7164b024333e6d02f8c6b2d017e4ecb.
Sampling: Python random.Random(42).sample(range(300), 50).
Rows are stored in random draw order. All seven original columns, their types,
values, and original IDs are preserved. No filtering or text changes were made.
The split is named train and contains exactly 50 rows. The… See the full description on the dataset page: https://huggingface.co/datasets/shihanlin/ambig-iac-sample.control_arena_iacmirror-ambig-iac
Ambig-IaC: Ambiguous Infrastructure-as-Code Benchmark
A benchmark dataset of 300 tasks for testing AI agents that generate Infrastructure-as-Code (Terraform) configurations from ambiguous natural language intents.
Project page: https://zyang37.github.io/ambig-iac.github.io/
Dataset Description
This dataset is sourced from IaC-Eval. We performed manual fixes to the original Terraform configurations and validated that all 300 tasks pass terraform plan. Each task… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-ambig-iac.iac-dataset
IAC Dataset
Intent-Aware Clustering benchmark datasets.
Each dataset is a separate config; select one with name= when loading:
from datasets import load_dataset
ds = load_dataset("gnitoahc/iac-dataset", name="arxiv")
Common Fields
Field
Type
Description
id
string
Original row index from the source dataset
text
string
Input text (may be multi-sentence or multi-line)
label
string
Class label for the text
Datasets… See the full description on the dataset page: https://huggingface.co/datasets/gnitoahc/iac-dataset.Futboljenny-tts-6h-taggedNeuralSpark-IACD-18kNeuralSpark-IACD-50k-v2jenny-tts-tags-6hNeuralSpark-IACD-V4-MASTER-138kNeuralSpark-IACD-MASTER-68kNeuralSpark-IACD-MASTER-138k
