datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ops-lite
ops-lite
A curated 500-case root-cause-analysis (RCA) evaluation set for
microservice systems, with manifest-driven causal-graph ground truth.
Each case bundles:
a chaos-injection ground truth (injection.json)
a causal service graph derived from the injection's fault contract
(causal_graph.json)
the runtime environment snapshot (env.json, result.json,
label.txt)
12 parquet metric tables per case, split into the abnormal window
(during fault) and the normal window (baseline)
The… See the full description on the dataset page: https://huggingface.co/datasets/anon-ops/ops-lite.git-ops-recovery-trajectories
Git Ops Recovery Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/git-ops-recovery-trajectories.op-spp-streams-v2
op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus)
Tokenized Dolma v1.7 (ODC-BY)
subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer +
<assistant> extension from epfl-dlab/spp-training):
compact dense-packed 2049-token windows; annotated/canary one document
per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run
(arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus
text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v2.Full_Agent_RL_OPSD_with_Just_2_A800sexcel-ai-ops
ExcelAI Ops Synthetic
Programmatic synthetic data for training a 1B specialist to generate Excel workbooks via constrained JSON ops.
Each row: instruction + input_data (CSV) -> completion (JSON string of ops list). Parse completion with json.loads, validate with schema.json, build with build_excel.py -> .xlsx.
Tasks: budget_tracker, invoice, sales_report, gradebook, inventory, timesheet.
Format:
prompt: "Instruction: ...\nData:\n...\nOutput JSON ops:"
completion: JSON string of… See the full description on the dataset page: https://huggingface.co/datasets/maxie-12321/excel-ai-ops.surgvu-cat2-vqa
SurgVU Category 2 VQA pairs
No video or frames are included. This is 23,354 question–answer pairs over
30-second windows of the published SurgVU dataset. Each record gives a case id and
a start/stop time, so anyone with the SurgVU videos can regenerate the exact frames.
Used to train the vision-language model in our SurgVU 2026 Category 2 submission.
Contents
file
data/train.jsonl
18,619 pairs
data/val.jsonl
4,735 pairs
recipe/build_qa_pairs.py… See the full description on the dataset page: https://huggingface.co/datasets/opscribe-ai/surgvu-cat2-vqa.tanpo-ops-sft
Tanpo Ops SFT (10k)
Ownership
Owner: DarkNinja Solutions
Creator: d4rkninja
Community: DarkLab
Domain
Operations and systems: process design, SOPs, tooling, capacity, quality, and scaling operational excellence.
Dataset summary
Field
Value
Rows
10000
Schema
Chat SFT (messages with system / user / assistant)
Source file
tanpo-ops-sft-10k-format-fixed.jsonl
Audit verdict
PASS
Empty assistant rate
0.0%
Exact… See the full description on the dataset page: https://huggingface.co/datasets/d4rkninja/tanpo-ops-sft.str-ops-corpus
STR-Ops-Corpus: Short-Term Rental Operations & Property Inspection Corpus
A domain-specific, openly licensed text corpus of 241 documents
(~275,747 words / ~435,376 tokens, cl100k_base) on short-term rental (STR) and vacation
rental property management: property inspection, turnover and changeover operations, cleaning,
damage detection and platform claims, maintenance, staffing, and regulatory compliance.
The corpus is intended for fine-tuning and evaluating large language… See the full description on the dataset page: https://huggingface.co/datasets/Metzpapa/str-ops-corpus.SPSD-Variants-opsd
SPSD-Variants-opsd
Grounded on-policy self-distillation (OPSD) teacher-context dataset over 45
board-game rule variants (5 families × 9: connect4, domineering,
simplified_first_attack, simplified_othello, tic_tac_chess), derived from
trained MuZero/EfficientZero checkpoints (plan-528 v2).
Each row is a decision-state task (a move choice or one of six auxiliary
state-QA tasks). The privileged_context is the teacher signal: grounded
natural-language reasoning that discovers the… See the full description on the dataset page: https://huggingface.co/datasets/LorMolf/SPSD-Variants-opsd.hr-ops-tools
HR-Ops: 8,621 rows of tool calling and cited policy for HR assistants
A training set for HR-operations assistants, built around one idea: make the HR task
objectively checkable. The headline shard is tool calling against authored HR-ops
function schemas, where a correct answer is exact JSON and a wrong one cannot hide behind
fluent prose. Built for the Adaption AutoScientist Challenge, Part 2 (HR).
What this dataset proves, and how you check it
rows
8… See the full description on the dataset page: https://huggingface.co/datasets/Jainamshahhh/hr-ops-tools.sn96g-security-opsec-2chunk1-20250919_190709
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 45.5 (median 45.5, min 39, max 52)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): 928cda462faf36edb78f69bf0ff0dcdc68ae2be35cf628878530dac670150a6a
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-security-opsec-2chunk1-20250919_190709.autonomous-revenue-ops
Autonomous Revenue Ops — Evaluation Dataset
Synthetic regression cases for validating deterministic revenue-operations policy behavior.
Dataset purpose
The dataset checks that structured qualification outputs map to the expected authorized workflow state. It is designed for regression testing, not for training a foundation model.
Covered behaviors
high-score, high-confidence autonomous routing
medium-score human review
low-confidence evidence… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/autonomous-revenue-ops.OPSD-PI-SWE-Gym-512
OPSD-PI SWE-Gym Stage PI 512
Qwen3.5-9B stage-adaptive OPSD-PI 的公开 512-row 数据与 Weak from-scratch
训练包。Public 512-row data and Weak from-scratch training bundle.
Files
data/train.jsonl: 512 deterministic SWE-Gym rows with Weak, Medium, and
Strong PI for EXPLORE, REPRODUCE, DIAGNOSE, EDIT, and VERIFY.
data/manifest.json: source selection and integrity metadata.
release/OPSD_pi-opsd-pi-weak-from-scratch-20260818.tar.gz: immutable
source release containing launchers… See the full description on the dataset page: https://huggingface.co/datasets/LSW142857/OPSD-PI-SWE-Gym-512.finance-ops-triage-v0.1
Finance Ops Triage v0.1 dataset
The original small, illustrative dataset prepared for Ugo Chukwu's first Unsloth fine-tuning and deployment exercise. The examples were provided during a guided ChatGPT experiment; they are not collected operational transaction records or an independently validated finance policy.
Code and experiment report · Model archive
Structure
Each JSONL row has messages containing system, user, and assistant entries. The assistant content is… See the full description on the dataset page: https://huggingface.co/datasets/Etherlabs/finance-ops-triage-v0.1.opsd-probe-seed
OPSD prefix-continuation probe — seed data
Everything needed to reproduce the prefix-continuation probe for OPSD (on-policy
self-distillation) on a fresh GPU box, except the base model (Qwen/Qwen3-1.7B, pulled from
the Hub at setup) and the code repo (hbin0701/OPSD).
These artifacts live outside git because the training/eval output directory is .gitignored.
What the probe answers
Fitting p' = p + λ·(1[mode correct] − p) + γ against a properly sampled 64-shot… See the full description on the dataset page: https://huggingface.co/datasets/hbin0701/opsd-probe-seed.op-spp-streams-v1
op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus)
Tokenized Dolma v1.7 (ODC-BY)
subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer +
<assistant> extension from epfl-dlab/spp-training):
compact dense-packed 2049-token windows; annotated/canary one document
per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run
(arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus
text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v1.replique-a
Description
The JSONL file generated by the script below contains detailed information about a corpus of public domain films, including their subtitles in multiple languages. Here is a detailed description of its structure:
JSONL file structure
IMDB: Unique identifier for the movie in the IMDb database.
primary_title: Primary title of the movie.
original_title: Original title of the movie.
french:
filepath: Relative path to the French subtitles file.
subtitles: List of… See the full description on the dataset page: https://huggingface.co/datasets/opsci/replique-a.opsd-plain-4b-rollouts
opsd-plain-4b-rollouts
This dataset contains rollout generations collected during training.
Source experiment
method: opsd-plain
model_size: 4b
experiment_dir: /home/irteam/outputs/opsd_plain_4b
Format
Each row contains:
step
sample_index
prompt
completion
method
model_size
source_file
Viewer structure
all: all rollout rows together
step_<N>: only one rollout step, easier to inspect in the dataset viewer
Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-4b-rollouts.call_of_duty_black_ops_iii_recordings_01
使命召唤12 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_a727c6fe84c92444c3fbf241e8380d26
Collection: general (泛数据)
Recordings: 7
Layout: recordings/<recording_id>/<raw component>
ru-quiz-qadevops-opsnotes-instructionsopsd-plain-8b-rollouts
opsd-plain-8b-rollouts
This dataset contains rollout generations collected during training.
Source experiment
method: opsd-plain
model_size: 8b
experiment_dir: /home/irteam/outputs/opsd_plain_8b
Format
Each row contains:
step
sample_index
prompt
completion
method
model_size
source_file
Viewer structure
all: all rollout rows together
step_<N>: only one rollout step, easier to inspect in the dataset viewer
Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-8b-rollouts.sn96g-security-opsec-10chunk1-20250919_134732
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 26
Avg answer length (tokens): 123.8 (median 125.0, min 91, max 159)
Schema errors: 0 (should be 0)
File size: 0.02 MB
SHA256 (data.jsonl): 00deb88f40dddfca0915bd9ea23c50f8fe74c9455b05f94d3e80ef446dfadd10
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-security-opsec-10chunk1-20250919_134732.ecommerce-ops-resultsAstreeAſtrée† is a repository of 2000 early modern instructions in French drawn from 162 French novels published between 1600 and 1700. Aſtrée can be used to fine-tuned any LLM on early modern French.
All the instructions have been created from one page excerpts extracting from public domain works with historical writing and typography. They may include OCR errors that should not affect significantly the quality of text generation.
Beyond their cultural relevance, Aſtrée provides a very good sample… See the full description on the dataset page: https://huggingface.co/datasets/opsci/Astree.eye-grep
eye-grep — log token-classification gold set
Token-level labels for eye-grep, a log colorizer that tags each content token of a
server-log line so a renderer can highlight ids, timestamps, IPs and repeated strings.
These are the sets used to train and evaluate the eye-grep taggers
(opsbr/eye-grep-deberta-v3-small
and the distilled opsbr/eye-grep-electra-small).
Fully synthetic and self-contained — generated by a deterministic template engine
(loop/synth_gold.py), with no… See the full description on the dataset page: https://huggingface.co/datasets/opsbr/eye-grep.ops_finetunesn96g-security-opsec-2chunk1-20250919_195736
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 35 (median 35.0, min 25, max 45)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): 7fcf981dc2ba25e93a757512227e7d695f56ff2400496c7454b41c2d10710c9f
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-security-opsec-2chunk1-20250919_195736.commerce-ops-results-3Paragraf_Opsiyon
