datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ops-lite
ops-lite
A curated 500-case root-cause-analysis (RCA) evaluation set for
microservice systems, with manifest-driven causal-graph ground truth.
Each case bundles:
a chaos-injection ground truth (injection.json)
a causal service graph derived from the injection's fault contract
(causal_graph.json)
the runtime environment snapshot (env.json, result.json,
label.txt)
12 parquet metric tables per case, split into the abnormal window
(during fault) and the normal window (baseline)
The… See the full description on the dataset page: https://huggingface.co/datasets/anon-ops/ops-lite.op-spp-streams-v2
op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus)
Tokenized Dolma v1.7 (ODC-BY)
subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer +
<assistant> extension from epfl-dlab/spp-training):
compact dense-packed 2049-token windows; annotated/canary one document
per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run
(arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus
text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v2.surgvu-cat2-vqa
SurgVU Category 2 VQA pairs
No video or frames are included. This is 23,354 question–answer pairs over
30-second windows of the published SurgVU dataset. Each record gives a case id and
a start/stop time, so anyone with the SurgVU videos can regenerate the exact frames.
Used to train the vision-language model in our SurgVU 2026 Category 2 submission.
Contents
file
data/train.jsonl
18,619 pairs
data/val.jsonl
4,735 pairs
recipe/build_qa_pairs.py… See the full description on the dataset page: https://huggingface.co/datasets/opscribe-ai/surgvu-cat2-vqa.opsd-probe-seed
OPSD prefix-continuation probe — seed data
Everything needed to reproduce the prefix-continuation probe for OPSD (on-policy
self-distillation) on a fresh GPU box, except the base model (Qwen/Qwen3-1.7B, pulled from
the Hub at setup) and the code repo (hbin0701/OPSD).
These artifacts live outside git because the training/eval output directory is .gitignored.
What the probe answers
Fitting p' = p + λ·(1[mode correct] − p) + γ against a properly sampled 64-shot… See the full description on the dataset page: https://huggingface.co/datasets/hbin0701/opsd-probe-seed.op-spp-streams-v1
op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus)
Tokenized Dolma v1.7 (ODC-BY)
subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer +
<assistant> extension from epfl-dlab/spp-training):
compact dense-packed 2049-token windows; annotated/canary one document
per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run
(arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus
text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v1.opsd-plain-4b-rollouts
opsd-plain-4b-rollouts
This dataset contains rollout generations collected during training.
Source experiment
method: opsd-plain
model_size: 4b
experiment_dir: /home/irteam/outputs/opsd_plain_4b
Format
Each row contains:
step
sample_index
prompt
completion
method
model_size
source_file
Viewer structure
all: all rollout rows together
step_<N>: only one rollout step, easier to inspect in the dataset viewer
Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-4b-rollouts.call_of_duty_black_ops_iii_recordings_01
使命召唤12 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_a727c6fe84c92444c3fbf241e8380d26
Collection: general (泛数据)
Recordings: 7
Layout: recordings/<recording_id>/<raw component>
opsd-plain-8b-rollouts
opsd-plain-8b-rollouts
This dataset contains rollout generations collected during training.
Source experiment
method: opsd-plain
model_size: 8b
experiment_dir: /home/irteam/outputs/opsd_plain_8b
Format
Each row contains:
step
sample_index
prompt
completion
method
model_size
source_file
Viewer structure
all: all rollout rows together
step_<N>: only one rollout step, easier to inspect in the dataset viewer
Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-8b-rollouts.ecommerce-ops-resultscommerce-ops-results-3commerce-ops-resultscommerce-ops-results-2commerce-ops-resultscommerce-ops-results-2
