CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anon-ops /ops-lite ops-lite A curated 500-case root-cause-analysis (RCA) evaluation set for microservice systems, with manifest-driven causal-graph ground truth. Each case bundles: a chaos-injection ground truth (injection.json) a causal service graph derived from the injection's fault contract (causal_graph.json) the runtime environment snapshot (env.json, result.json, label.txt) 12 parquet metric tables per case, split into the abnormal window (during fault) and the normal window (baseline) The… See the full description on the dataset page: https://huggingface.co/datasets/anon-ops/ops-lite.tabulargraph-mln<1K1 likes1k downloads4mo agoHugging Face02joshycodes /op-spp-streams-v2 op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus) Tokenized Dolma v1.7 (ODC-BY) subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer + <assistant> extension from epfl-dlab/spp-training): compact dense-packed 2049-token windows; annotated/canary one document per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run (arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v2.tabularn<1K0 likes94 downloads1mo agoHugging Face03opscribe-ai /surgvu-cat2-vqa SurgVU Category 2 VQA pairs No video or frames are included. This is 23,354 question–answer pairs over 30-second windows of the published SurgVU dataset. Each record gives a case id and a start/stop time, so anyone with the SurgVU videos can regenerate the exact frames. Used to train the vision-language model in our SurgVU 2026 Category 2 submission. Contents file data/train.jsonl 18,619 pairs data/val.jsonl 4,735 pairs recipe/build_qa_pairs.py… See the full description on the dataset page: https://huggingface.co/datasets/opscribe-ai/surgvu-cat2-vqa.tabularvisual-question-answering10K<n<100K0 likes73 downloads9d agoHugging Face04hbin0701 /opsd-probe-seed OPSD prefix-continuation probe — seed data Everything needed to reproduce the prefix-continuation probe for OPSD (on-policy self-distillation) on a fresh GPU box, except the base model (Qwen/Qwen3-1.7B, pulled from the Hub at setup) and the code repo (hbin0701/OPSD). These artifacts live outside git because the training/eval output directory is .gitignored. What the probe answers Fitting p' = p + λ·(1[mode correct] − p) + γ against a properly sampled 64-shot… See the full description on the dataset page: https://huggingface.co/datasets/hbin0701/opsd-probe-seed.tabulartext-generationn<1K0 likes36 downloads1mo agoHugging Face05joshycodes /op-spp-streams-v1 op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus) Tokenized Dolma v1.7 (ODC-BY) subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer + <assistant> extension from epfl-dlab/spp-training): compact dense-packed 2049-token windows; annotated/canary one document per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run (arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v1.tabularn<1K0 likes35 downloads1mo agoHugging Face06SeongryongJung /opsd-plain-4b-rollouts opsd-plain-4b-rollouts This dataset contains rollout generations collected during training. Source experiment method: opsd-plain model_size: 4b experiment_dir: /home/irteam/outputs/opsd_plain_4b Format Each row contains: step sample_index prompt completion method model_size source_file Viewer structure all: all rollout rows together step_<N>: only one rollout step, easier to inspect in the dataset viewer Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-4b-rollouts.tabulartext-generationn<1K0 likes26 downloads4mo agoHugging Face07opensima24 /call_of_duty_black_ops_iii_recordings_01gated 使命召唤12 raw recordings This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval. Game ID: game_a727c6fe84c92444c3fbf241e8380d26 Collection: general (泛数据) Recordings: 7 Layout: recordings/<recording_id>/<raw component> tabularn<1K0 likes24 downloads25d agoHugging Face08SeongryongJung /opsd-plain-8b-rollouts opsd-plain-8b-rollouts This dataset contains rollout generations collected during training. Source experiment method: opsd-plain model_size: 8b experiment_dir: /home/irteam/outputs/opsd_plain_8b Format Each row contains: step sample_index prompt completion method model_size source_file Viewer structure all: all rollout rows together step_<N>: only one rollout step, easier to inspect in the dataset viewer Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-8b-rollouts.tabulartext-generationn<1K0 likes16 downloads4mo agoHugging Face09Gloomytarsier3 /ecommerce-ops-resultstabularn<1K0 likes14 downloads5mo agoHugging Face10TenduL /commerce-ops-results-3tabularn<1K0 likes7 downloads5mo agoHugging Face11TenduL /commerce-ops-resultstabularn<1K0 likes6 downloads5mo agoHugging Face12TenduL /commerce-ops-results-2tabularn<1K0 likes3 downloads5mo agoHugging Face13Gloomytarsier3 /commerce-ops-resultstabularn<1K0 likes3 downloads5mo agoHugging Face14Gloomytarsier3 /commerce-ops-results-2tabularn<1K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.