CoolFace
Modelpublic

Arijit-07/aria-devops-llama8b

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes
Model Card

ARIA — DevOps Incident Response Agent

Llama-3.1-8B fine-tuned with GRPO

Trained on the ARIA DevOps Incident Response live RL environment using GRPO.

Training Results

TaskBaselineFine-tunedImprovement
easy0.3200.685+0.365
medium0.0500.378+0.328
hard0.1900.869+0.679
bonus0.1520.682+0.530

[image]

Setup

  • Algorithm: GRPO
  • Base: Llama-3.1-8B-Instruct
  • LoRA rank: 32, alpha: 64
  • Episodes: 160 (40 per task)
  • GPU: NVIDIA L4, 162 minutes
  • Framework: Unsloth + HuggingFace TRL

Links

  • Environment: https://huggingface.co/spaces/Arijit-07/devops-incident-response
  • GitHub: https://github.com/Twilight-13/devops-incident-response