Arijit-07/aria-devops-llama8b
0
ARIA — DevOps Incident Response Agent
Llama-3.1-8B fine-tuned with GRPO
Trained on the ARIA DevOps Incident Response live RL environment using GRPO.
Training Results
Setup
- Algorithm: GRPO
- Base: Llama-3.1-8B-Instruct
- LoRA rank: 32, alpha: 64
- Episodes: 160 (40 per task)
- GPU: NVIDIA L4, 162 minutes
- Framework: Unsloth + HuggingFace TRL
Links
- Environment: https://huggingface.co/spaces/Arijit-07/devops-incident-response
- GitHub: https://github.com/Twilight-13/devops-incident-response
