CoolFace
Apppublic

Abiraminayagi/bug-triage-rl

sourceHugging Faceupdated 5mo agoView on Hugging Face
2likes
App README

Bug Triage & Escalation Desk — OpenEnv RL Environment

A high-fidelity simulation environment for AI-driven DevOps decision intelligence. The system models dynamic bug queues, SLA constraints, and resource-limited developer teams, creating a realistic operational setting for evaluating autonomous agents.

Unlike toy benchmarks, this environment captures real-world trade-offs and uncertainty, serving as a training and evaluation ground for next-generation AI agents in software engineering workflows.

Motivation

Modern software systems deal with massive volumes of bugs, incidents, and operational tasks daily. Efficient triage requires balancing SLA deadlines, developer workload, task severity, and system stability simultaneously.

This environment simulates this complex decision-making process, enabling the development and evaluation of intelligent agents for DevOps automation — exactly what engineering managers at companies like Meta do every day.

Environment Overview

PropertyValue
FrameworkOpenEnv + FastAPI
Action SpaceDiscrete(4): assign, escalate, defer, close
Observation SpaceDict: bug queue + team state + system metrics
Reward Range[-1.0, 1.0]
Max Episode Steps100
Taskseasy, medium, hard

Environment Design

The environment simulates a dynamic operational system with:

  • —Dynamic task queue — incoming bugs with varying severity (Low, Medium, High, Critical)
  • —SLA constraints — deadlines and penalties for overdue bugs
  • —Resource-constrained team — developers with limited capacity and different skills
  • —Evolving system state — queue health, workload, backlog tracking

Each decision impacts long-term system performance and is rewarded accordingly.

Action Space

ActionDescriptionGood When
assignAssign bug to available developerDeveloper available with matching skills
escalateEscalate to urgent statusBug is overdue (SLA > 100%) or Critical severity
deferPush bug to laterLow priority bug, team fully loaded
closeMark bug as resolvedBug is fixed or invalid

Observation Space

json
{
  "bug_queue": {
    "open_bugs": [
      {
        "id": "BUG-001",
        "title": "Authentication failing on mobile",
        "severity": "critical",
        "bug_type": "backend",
        "age_hours": 6,
        "affected_users": 1250,
        "priority_score": 0.95,
        "sla_hours": 4,
        "sla_usage_pct": 150.0,
        "is_overdue": true
      }
    ],
    "total_count": 15,
    "open_count": 12,
    "critical_count": 2,
    "queue_health_score": 0.73
  },
  "team": {
    "developers": [
      {
        "id": "dev-001",
        "name": "Sarah Chen",
        "skills": ["frontend", "backend"],
        "current_load": 2,
        "max_capacity": 3,
        "is_available": true
      }
    ],
    "availability_ratio": 0.67
  },
  "metrics": {
    "step": 5,
    "max_steps": 100,
    "cumulative_reward": 1.85,
    "task_level": "medium"
  }
}

Reward Function

ComponentDescription
Base rewardAction type baseline (+0.2 to +0.5)
Severity modifierCritical bugs handled correctly = big bonus
SLA rewardAddressing overdue bugs = positive reward
Efficiency rewardAssigning to least-loaded available developer
Queue healthImprovement in overall queue health score
PenaltiesDeferring critical bugs, unnecessary escalations

Agent

The environment is evaluated using Qwen 72B via HuggingFace router, acting as an autonomous decision-making agent. The agent interprets the bug queue state, reasons about SLA constraints and developer capacity, and selects optimal triage actions. This demonstrates the potential of LLMs in structured DevOps decision-making beyond text generation.

Tasks

Easy

  • —8 bugs (Low/Medium severity only)
  • —3 developers, all available
  • —Relaxed SLA deadlines (48h+)
  • —Passing score: 0.5
  • —Baseline score: 0.95 - 0.99

Medium

  • —15 bugs (mixed severity, some Critical)
  • —4 developers, partial availability
  • —Moderate SLA pressure (12-24h)
  • —Passing score: 0.6
  • —Baseline score: 0.85 - 0.92

Hard

  • —25 bugs (mostly Critical/High)
  • —8 developers but overwhelmed
  • —Most SLAs already overdue
  • —New bugs stream in every 10 steps
  • —Passing score: 0.7
  • —Baseline score: 0.40 - 0.80

Results

Task LevelScore RangeStatus
Easy0.95 - 0.99PASS
Medium0.85 - 0.92PASS
Hard0.40 - 0.80PASS
Average0.73 - 0.90PASS

The agent performs strongly in structured scenarios and demonstrates robustness under increasing complexity.

Why This Matters

This environment represents a step toward AI-driven DevOps automation — intelligent task prioritization at scale. Every software company from startups to Meta deals with bug triage daily. This benchmark enables evaluating AI agents in real engineering workflows, providing a foundation for next-generation autonomous software operations.

API Endpoints

EndpointMethodDescription
/GETEnvironment info
/healthGETHealth check
/resetPOSTStart new episode
/stepPOSTTake one action
/stateGETCurrent episode state
/gradeGETEpisode score (0.0-1.0)
/tasksGETList all tasks
/docsGETInteractive API docs

Quick Start

bash
git clone https://github.com/Abirami-2743/bug-triage-rl
cd bug-triage-rl
docker build -t bug-triage-rl .
docker run -p 7860:7860 bug-triage-rl
curl http://localhost:7860/health

Run Inference

bash
export API_BASE_URL="https://router.huggingface.co/v1"
export MODEL_NAME="Qwen/Qwen2.5-72B-Instruct"
export HF_TOKEN="your-hf-token"
python inference.py

Project Structure

bug-triage-rl/
├── server/
│   ├── app.py                    # FastAPI server
│   ├── bug_triage_environment.py # Core RL environment
│   └── __init__.py
├── src/
│   ├── models.py                 # Pydantic models
│   ├── bug_generator.py          # Bug generation
│   ├── reward_function.py        # Reward calculation
│   └── environment_gymnasium.py  # Gymnasium wrapper
├── inference.py                  # Baseline inference script
├── Dockerfile                    # Container definition
├── openenv.yaml                  # OpenEnv manifest
├── pyproject.toml                # Package config
├── requirements.txt              # Dependencies
└── README.md

Live Demo

  • —HF Space: https://huggingface.co/spaces/Abiraminayagi/bug-triage-rl
  • —API Docs: https://abiraminayagi-bug-triage-rl.hf.space/docs
  • —Health: https://abiraminayagi-bug-triage-rl.hf.space/health

Built For

OpenEnv AI Hackathon 2026 - Meta x Hugging Face x Scaler School of Technology