CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01siyanzhao /Openthoughts_math_30k_opsdtext10K<n<100K10 likes13k downloads7mo agoHugging Face02anon-ops /ops-lite ops-lite A curated 500-case root-cause-analysis (RCA) evaluation set for microservice systems, with manifest-driven causal-graph ground truth. Each case bundles: a chaos-injection ground truth (injection.json) a causal service graph derived from the injection's fault contract (causal_graph.json) the runtime environment snapshot (env.json, result.json, label.txt) 12 parquet metric tables per case, split into the abnormal window (during fault) and the normal window (baseline) The… See the full description on the dataset page: https://huggingface.co/datasets/anon-ops/ops-lite.tabulargraph-mln<1K1 likes1k downloads4mo agoHugging Face03BytedTsinghua-SIA /CUDA-Agent-Ops-6K CUDA-Agent-Ops-6K CUDA-Agent-Ops-6K is a curated training dataset for CUDA kernel generation and optimization. It is released as part of the CUDA-Agent project: Project Page: https://CUDA-Agent.github.io/ Github Repo: https://github.com/BytedTsinghua-SIA/CUDA-Agent Dataset Summary CUDA-Agent-Ops-6K contains 6,000 synthesized operator-level training tasks designed for large-scale agentic RL training. It is intended to provide diverse and executable CUDA-oriented training… See the full description on the dataset page: https://huggingface.co/datasets/BytedTsinghua-SIA/CUDA-Agent-Ops-6K.texttext-generation1K<n<10K92 likes555 downloads7mo agoHugging Face04yunjae-won /openthoughts_math_30k_opsd_splittext10K<n<100K0 likes462 downloads3mo agoHugging Face05rmems /git-ops-recovery-trajectories Git Ops Recovery Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/git-ops-recovery-trajectories.text1K<n<10K0 likes452 downloads22d agoHugging Face06abhijitbetigeri /dc-ops-dataset DC-Ops: Data Center Components Dataset On-device data center operations assistant dataset for the Qualcomm x Meta ExecuTorch Hackathon. Overview 319 images of data center infrastructure (server racks, NVIDIA NVL72, cables, ports, etc.) 3,045 polygon annotations in YOLO-seg format 16 component classes auto-labeled with Grounding DINO + SAM, for fine-tuning YOLOv8n-seg Classes ID Class Count 0 server rack rack enclosures 1 compute tray… See the full description on the dataset page: https://huggingface.co/datasets/abhijitbetigeri/dc-ops-dataset.imageobject-detection1K<n<10K0 likes435 downloads3mo agoHugging Face07ardauzunoglu /bitwise-opstabular100K<n<1M0 likes254 downloads7mo agoHugging Face08prestonfu /gsm_infinite_hard_r0.4_ops24text10K<n<100K0 likes219 downloads5mo agoHugging Face09ardauzunoglu /string-opstext100K<n<1M0 likes207 downloads6mo agoHugging Face10HzChen20 /opsd-cl-assets opsd-cl-assets Offline assets for the OPSD continual-learning project on a box without DNS. Built 2026-09-12 on VAST. Pull everything with one command, then run install_on_box.sh. path content size wheelhouse/ 167 wheels resolved from siyan-zhao/OPSD environment.yml for cp310 / manylinux x86_64, plus deepspeed-0.18.2.tar.gz and the lock file 4.7 GB models/Qwen3-1.7B Qwen/Qwen3-1.7B 3.8 GB datasets/Openthoughts_math_30k_opsd siyanzhao/Openthoughts_math_30k_opsd… See the full description on the dataset page: https://huggingface.co/datasets/HzChen20/opsd-cl-assets.text0 likes188 downloads8d agoHugging Face11anon-ops /ops-lite-review-sample ops-lite-review-sample This repository is a reviewer-facing sample of the full ops-lite benchmark. Full dataset: https://huggingface.co/datasets/anon-ops/ops-lite Full dataset release used for this sample: the same public ops-lite release associated with this submission Sample dataset URL: https://huggingface.co/datasets/anon-ops/ops-lite-review-sample What is included 10 complete RCA cases under cases/ a subset manifest.jsonl containing metadata for those same 10… See the full description on the dataset page: https://huggingface.co/datasets/anon-ops/ops-lite-review-sample.textother1M<n<10M0 likes171 downloads5mo agoHugging Face12joshycodes /op-spp-streams-v2 op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus) Tokenized Dolma v1.7 (ODC-BY) subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer + <assistant> extension from epfl-dlab/spp-training): compact dense-packed 2049-token windows; annotated/canary one document per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run (arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v2.tabularn<1K0 likes94 downloads1mo agoHugging Face13Camellia86 /Full_Agent_RL_OPSD_with_Just_2_A800stext100K<n<1M0 likes90 downloads23d agoHugging Face14maxie-12321 /excel-ai-ops ExcelAI Ops Synthetic Programmatic synthetic data for training a 1B specialist to generate Excel workbooks via constrained JSON ops. Each row: instruction + input_data (CSV) -> completion (JSON string of ops list). Parse completion with json.loads, validate with schema.json, build with build_excel.py -> .xlsx. Tasks: budget_tracker, invoice, sales_report, gradebook, inventory, timesheet. Format: prompt: "Instruction: ...\nData:\n...\nOutput JSON ops:" completion: JSON string of… See the full description on the dataset page: https://huggingface.co/datasets/maxie-12321/excel-ai-ops.text10K<n<100K0 likes80 downloads17d agoHugging Face15prestonfu /gsm_infinite_hard_r0.4_ops16text10K<n<100K0 likes78 downloads5mo agoHugging Face16tarekmasryo /llm-system-ops-production-telemetry-sft-data 🤖📈 LLM System Ops Telemetry (Synthetic) A synthetic, production-style, multi-table LLM telemetry dataset designed for LLMOps analytics and decision-grade experiments. It supports monitoring cost, latency, tokens, failures, safety flags, tool usage, and user feedback at the interaction level, with rollups at the session and user levels — plus an SFT table aligned 1:1 with interactions and a prompt/config dimension. Synthetic data (safe for teaching, prototyping, and portfolio… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/llm-system-ops-production-telemetry-sft-data.tabulartabular-classification10K<n<100K1 likes76 downloads8mo agoHugging Face17starli-snowflake /swe_opsd_datasettext10K<n<100K0 likes76 downloads3mo agoHugging Face18locaszzz /opsd-tooltext10K<n<100K0 likes76 downloads6d agoHugging Face19opscribe-ai /surgvu-cat2-vqa SurgVU Category 2 VQA pairs No video or frames are included. This is 23,354 question–answer pairs over 30-second windows of the published SurgVU dataset. Each record gives a case id and a start/stop time, so anyone with the SurgVU videos can regenerate the exact frames. Used to train the vision-language model in our SurgVU 2026 Category 2 submission. Contents file data/train.jsonl 18,619 pairs data/val.jsonl 4,735 pairs recipe/build_qa_pairs.py… See the full description on the dataset page: https://huggingface.co/datasets/opscribe-ai/surgvu-cat2-vqa.tabularvisual-question-answering10K<n<100K0 likes73 downloads8d agoHugging Face20Opscidia /patentstext10K<n<100K0 likes72 downloads2y agoHugging Face21LHL3341 /llava-10k-cot-iqa-opsimage10K<n<100K0 likes72 downloads1y agoHugging Face22ram-lexsi /agenttune-agent-ops-SFT-qwen3-0.6b-tracestabularn<1K0 likes69 downloads11d agoHugging Face23ram-lexsi /agenttune-agent-ops-GRPO-tracestabularn<1K0 likes68 downloads12d agoHugging Face24prestonfu /gsm_infinite_hard_r0.4_ops8text10K<n<100K0 likes67 downloads5mo agoHugging Face25d4rkninja /tanpo-ops-sft Tanpo Ops SFT (10k) Ownership Owner: DarkNinja Solutions Creator: d4rkninja Community: DarkLab Domain Operations and systems: process design, SOPs, tooling, capacity, quality, and scaling operational excellence. Dataset summary Field Value Rows 10000 Schema Chat SFT (messages with system / user / assistant) Source file tanpo-ops-sft-10k-format-fixed.jsonl Audit verdict PASS Empty assistant rate 0.0% Exact… See the full description on the dataset page: https://huggingface.co/datasets/d4rkninja/tanpo-ops-sft.texttext-generation10K<n<100K0 likes65 downloads5d agoHugging Face26Metzpapa /str-ops-corpus STR-Ops-Corpus: Short-Term Rental Operations & Property Inspection Corpus A domain-specific, openly licensed text corpus of 241 documents (~275,747 words / ~435,376 tokens, cl100k_base) on short-term rental (STR) and vacation rental property management: property inspection, turnover and changeover operations, cleaning, damage detection and platform claims, maintenance, staffing, and regulatory compliance. The corpus is intended for fine-tuning and evaluating large language… See the full description on the dataset page: https://huggingface.co/datasets/Metzpapa/str-ops-corpus.texttext-generationn<1K0 likes61 downloads3mo agoHugging Face27LorMolf /SPSD-Variants-opsd SPSD-Variants-opsd Grounded on-policy self-distillation (OPSD) teacher-context dataset over 45 board-game rule variants (5 families × 9: connect4, domineering, simplified_first_attack, simplified_othello, tic_tac_chess), derived from trained MuZero/EfficientZero checkpoints (plan-528 v2). Each row is a decision-state task (a move choice or one of six auxiliary state-QA tasks). The privileged_context is the teacher signal: grounded natural-language reasoning that discovers the… See the full description on the dataset page: https://huggingface.co/datasets/LorMolf/SPSD-Variants-opsd.texttext-generation100K<n<1M0 likes59 downloads23d agoHugging Face28Jainamshahhh /hr-ops-tools HR-Ops: 8,621 rows of tool calling and cited policy for HR assistants A training set for HR-operations assistants, built around one idea: make the HR task objectively checkable. The headline shard is tool calling against authored HR-ops function schemas, where a correct answer is exact JSON and a wrong one cannot hide behind fluent prose. Built for the Adaption AutoScientist Challenge, Part 2 (HR). What this dataset proves, and how you check it rows 8… See the full description on the dataset page: https://huggingface.co/datasets/Jainamshahhh/hr-ops-tools.texttext-generation1K<n<10K0 likes55 downloads1mo agoHugging Face29starli-snowflake /scaleswe-opsd-v2-3200-summary Scale-SWE OPSD v2 — 3200 tasks with summary hints The training set used for the Scale-SWE on-policy self-distillation (OPSD) runs. 3200 SWE tasks across 752 repositories, each paired with a reference agent trajectory and a condensed solution hint. Uploaded from /checkpoint/huggingface/datasets/scaleswe_opsd_v2_3200_summary (a datasets.save_to_disk directory), converted to parquet. Row count, ids and field contents verified identical to the source. ⚠️ Contains… See the full description on the dataset page: https://huggingface.co/datasets/starli-snowflake/scaleswe-opsd-v2-3200-summary.texttext-generation1K<n<10K0 likes54 downloads2mo agoHugging Face30mt0rm0 /opsdataDataset Card for Energy — OPEN POWER System Data This dataset was prepared for the OpenHPI course Time Series Analysis and Forecasting and provides electricity consumption, renewable generation, weather, and electricity price data across European countries, including Germany. The data is aggregated from the ENTSO-E Transparency Platform and the Open Power System Data project, spanning over 10 years with 15- and 30-minute time intervals. It is suitable for energy forecasting, renewable… See the full description on the dataset page: https://huggingface.co/datasets/mt0rm0/opsdata.text10M<n<100M1 likes53 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.