arun-misra/ai-soar-training
0
AI-SOAR GRPO Training Space
Trains Qwen2.5-7B-Instruct with LoRA using GRPO reinforcement learning on the Network Defense environment.
Setup
- Set Secrets in Space settings:
HF_TOKEN— HuggingFace token with write accessWANDB_API_KEY— your wandb API key
- Hardware: A10G small (24 GB VRAM) minimum
Files needed in this Space repo
app.py
requirements.txt
models.py ← copy from vir_env/
server/ ← copy from vir_env/server/
__init__.py
vir_env_environment.py
app.py