CoolFace
Apppublic

Nani4481/datacenter-thermal-env

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes
App README

๐ŸŒก๏ธ Hyperscale AI Data Center Thermal & Power Controller

Solving the Physical Bottleneck of AI Progress

![License](https://opensource.org/licenses/Apache-2.0) ![Python](https://www.python.org/) ![Docker](https://www.docker.com/) ![FastAPI](https://fastapi.tiangolo.com/) ![OpenEnv](https://openenv.ai)


๐Ÿ“– Overview

In the race to train the next generation of Frontier Models โ€” like Llama-4 or Gemini-3 โ€” the absolute biggest physical bottleneck isn't silicon. It's power and cooling.

Massive GPU clusters consume megawatts of electricity. If GPUs overheat, they undergo thermal throttling, killing training performance โ€” or worse, triggering hardware failure worth millions of dollars.

This project provides a high-fidelity OpenEnv environment that simulates a 2D grid of server racks. It challenges AI agents to act as Thermal & Power Controllers, balancing computational throughput with energy-efficient cooling while navigating hardware failures โ€” a real problem at the frontier of AI infrastructure.


๐Ÿ—๏ธ System Architecture

The environment models a complex, stateful thermodynamic system with realistic physics. To ensure stability during concurrent automated evaluations, the API backend implements a robust Least Recently Used (LRU) Cache for session memory management.

Physics LayerDescription
Heat GenerationProportional to GPU workload (0.4ยฐC per % load)
Cooling PhysicsNon-linear HVAC efficiency โ€” cooling power scales cubically with output
Thermal BleedHeat diffuses between adjacent racks via neighbor-state differential
Ambient DriftNatural heating/cooling based on data center ambient temperature
text
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚              DATA CENTER SIMULATION GRID                โ”‚
โ”‚                                                         โ”‚
โ”‚   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”โ”‚
โ”‚   โ”‚  RACK R0 โ”‚  โ”‚  RACK R1 โ”‚  โ”‚  RACK R2 โ”‚  โ”‚  RACK R3 โ”‚โ”‚
โ”‚   โ”‚ ๐Ÿ”ฅ 85ยฐC  โ”‚โ”€โ”€โ”‚ ๐ŸŒก๏ธ 72ยฐC  โ”‚โ”€โ”€โ”‚ โœ… 60ยฐC  โ”‚โ”€โ”€โ”‚ โœ… 58ยฐC  โ”‚โ”‚
โ”‚   โ”‚  Load:90%โ”‚  โ”‚  Load:65%โ”‚  โ”‚  Load:40%โ”‚  โ”‚  Load:35%โ”‚โ”‚
โ”‚   โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜โ”‚
โ”‚        โ”‚             โ”‚             โ”‚             โ”‚      โ”‚
โ”‚   โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚
โ”‚   โ”‚          HVAC ZONE H0         HVAC ZONE H1          โ”‚ โ”‚
โ”‚   โ”‚          Output: 78%              Output: 45%       โ”‚ โ”‚
โ”‚   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿ› ๏ธ OpenEnv Specification

1. Observation Space โ€” DataCenterObservation

The agent receives a complete telemetry snapshot of the data center at every timestep, strictly typed via Pydantic:

python
class DataCenterObservation(BaseModel):
    step: int
    racks: List[RackState]          # Per-rack telemetry (temp, load, status)
    hvacs: List[HVACState]          # HVAC zone status (output %, operational status)
    current_pue: float              # Power Usage Effectiveness
    total_compute_kw: float
    total_cooling_kw: float
    load_imbalance: float           # Max-rack load โˆ’ Min-rack load (%)
    thermal_warnings: int           # Racks exceeding 80ยฐC

2. Action Space โ€” DataCenterAction

The agent controls the physical layer through three primary vectors:

python
class DataCenterAction(BaseModel):
    hvac_adjustments: Dict[str, float]   # Map of hvac_id to new cooling_output_percent (0-100)
    workload_shifts: List[WorkloadShift] # Shift compute load: {from_rack, to_rack, amount_percent}
    throttles: Dict[str, float]          # Map of rack_id to throttle amount (reduce workload)
โš ๏ธ throttles is a destructive action. Reducing workload saves hardware but sacrifices compute SLA. The agent must mathematically deduce when safe racks lack the capacity to absorb failing workloads and trigger throttles to survive.

3. Reward Function (Strictly Bounded)

The reward signal is shaped to incentivize efficiency while penalizing risk. Per hackathon rules, all rewards and final grades are strictly clamped within the bounds of [0.01, 0.99].

$$R = \text{Clamp}_{0.01}^{0.99} \left( f(\text{PUE}) - 0.1 \times |\text{Warnings}| - 0.5 \times |\text{Violations}| \right)$$

SignalConditionImpact
PUE BonusPUE โ‰ค 1.25Positive scaling
Warning PenaltyRack temp 80ยฐC โ€“ 89ยฐC-0.1 per rack
Violation PenaltyRack temp โ‰ฅ 90ยฐC-0.5 per rack

๐ŸŽฏ Tasks & Difficulty

#Task NameObjectiveCore Challenge
๐ŸŸข EasySteady StateAchieve PUE โ‰ค 1.25Balance cooling vs. power in a static environment without triggering thermal warnings.
๐ŸŸก MediumLoad SurgeHandle 100% load spike on R0Rapid workload migration (shifts) to idle racks before thermal violation occurs.
๐Ÿ”ด HardHVAC FailureH0 fails; evacuate R0/R1Capacity Crisis: Safe racks lack full capacity to absorb the failing workload. The agent MUST logically deduce the need to use destructive throttles alongside shifts to prevent a total meltdown.

๐Ÿš€ Getting Started

Prerequisites

  • โ€”Docker installed and running
  • โ€”Python 3.12+
  • โ€”Hugging Face Token (HF_TOKEN): Required to authenticate via the Hugging Face inference router.

๐Ÿณ Local Development

Step 1 โ€” Build the container:

bash
docker build -t datacenter-env .

Step 2 โ€” Run the FastAPI simulation server:

bash
docker run -p 7860:7860 datacenter-env
The server implements an LRU Cache to safely handle up to 50 concurrent validation sessions without memory leaks.

Step 3 โ€” Run the inference agent (in a new terminal):

bash
export HF_TOKEN="hf_your_token_here"
python inference.py

๐Ÿ“Š Baseline Results

Evaluated using Frontier-class reasoning models (Qwen-2.5-72B-Instruct / Llama-3.3-70B):

TaskDifficultyResultFinal Score
Steady State๐ŸŸข Easyโœ… Success0.99
Load Surge๐ŸŸก Mediumโœ… Success0.99
HVAC Failure๐Ÿ”ด Hardโœ… Success0.99
Note: High-reasoning models successfully utilize the throttles fallback action in the Hard task, proving the environment is legitimately solvable by LLMs reacting to state telemetry.

๐Ÿ“ Project Structure

plaintext
datacenter-env/
โ”œโ”€โ”€ Dockerfile               # Container definition & healthchecks
โ”œโ”€โ”€ openenv.yaml             # Standardized OpenEnv task configuration
โ”œโ”€โ”€ datacenter_env.py        # Core thermodynamics engine and strict Pydantic schemas
โ”œโ”€โ”€ inference.py             # Agent entry point with structured prompt engineering
โ””โ”€โ”€ server/
    โ””โ”€โ”€ app.py               # FastAPI server with memory-safe LRU caching

๐Ÿ“œ License

This project is licensed under the Apache License 2.0.


<div align="center">

Created for the Meta x Scaler OpenEnv 2026 Challenge ๐Ÿ†

Pushing the physical limits of AI infrastructure, one thermal cycle at a time.

</div>