datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datacenter-thermal-performance-coherence-risk-v0.1What this repo is for
Detect when heat and performance start to decouple.
Focus
• hotspot temperature vs throttling
• cooling margin vs p95 latency stability
• fan headroom vs sustained load
• early drift before alarms and shutdowns
Why it matters
Many outages are preceded by thermal drift.
Operators see it as “random latency.”
This dataset turns it into an early warning signal.
datacenter-network-routing-coherence-risk-v0.1What this repo is for
Detect routing instability before outages.
Focus
• packet loss vs route flapping
• jitter vs path stability
• link utilization vs reroutes
• early latency drift
Why it matters
Most outages show network symptoms before full failure.
This dataset turns those symptoms into early warnings.
datacenter-node-failure-cascade-risk-v0.1What this repo is for
Detect cascading node failure risk before full cluster outage.
Focus
• replication margin vs load
• failover speed
• service dependency density
• node failure clustering
Why it matters
Most major outages begin with small local failures.
The cascade forms when redundancy, load shift, and recovery fall out of alignment.
This dataset lets operators detect that drift early.
datacenter-power-load-coherence-risk-v0.1What this repo is for
Detect power instability before outages.
Focus
• rack load vs PSU margin
• UPS buffer vs spike frequency
• voltage variance vs reset events
• sustained load vs safe envelope
Why it matters
Power issues often look random.
They usually follow slow margin erosion.
This dataset turns that erosion into a visible signal.
datacenter-water-cooling-demand-coherence-risk-v0.1What this repo is for
Detect when datacenter compute heat
outpaces cooling water capacity.
Flags
drought stress and discharge limits under high heat load
curtailment events under rising demand
heat and water draw decouple from throttling
abnormal overdraw at low heat load
datacenter-job-queue-resource-coherence-risk-v0.1What this repo is for
Detect scheduler pressure before customer-visible incidents.
Focus
• queue depth vs wait time
• utilization vs SLO breaches
• preemption vs job failure
• fragmentation where capacity exists but cannot be used
Why it matters
Operators often see queue problems too late.
This dataset spots the drift earlier and links it to real failure signals.
