lokeshrao226/eco-logistics-qwen-grpo-v2
Eco-Logistics Qwen-2.5-1.5B GRPO (v9)
GRPO-trained LoRA adapter for Qwen-2.5-1.5B-Instruct that plans 25-step supply-chain trajectories under a strict carbon budget.
Submitted to OpenEnv Hackathon Round 2 (Crystal Blue) — World Modeling for Professional Tasks.
Headline Numbers
Held-out eval on 10 seeds × 3 tasks = 30 episodes, greedy decoding, with constrained parser:
Architecture
- Base:
Qwen/Qwen2.5-1.5B-Instruct, 4-bit quantized - Adapter: LoRA r=16, target modules
[q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj] - Trainable parameters: 18.5M (1.18% of base)
Training
- Method: GRPO via TRL 0.24, Unsloth wrapping, single T4 GPU
- Reward: grader-only (avoids the SFT-then-GRPO collapse seen in v8 experiments)
- Steps: 40 GRPO steps, ~58 min on T4
- Hyperparameters: 4 generations/prompt, batch 4, grad_accum 4, lr 1.5e-6
- Prompts: 80 unique initial states (seeds 0–119)
- Task mix during training: 60% netzeroprofit, 40% inventory_balanced
Inference (CRITICAL: use constrained parser)
The model is paired with a constrained parser that projects model output to feasible actions:
def parse_plan(completion, target_length=25):
# ... extract JSON array ...
cleaned.append({
"ship_amount": min(2.0, max(0.0, float(item.get("ship_amount", 0.0)))),
"origin_city": _coerce_city(item.get("origin_city", ""), "Seattle"),
"destination_city": _coerce_city(item.get("destination_city", ""), "Chicago"),
"speed_mode": "Rail", # FORCED — Air carbon kills the budget
})The cap ship≤2 is derived from carbon-budget arithmetic: 25 steps × 2 ship × 2 rail-carbon ≤ 100 budget.
Without the parser, raw model output achieves grader 0.144 on net_zero_profit. With the parser, 0.338. Treat the LoRA + parser as a single policy.
Quick Use
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct", load_in_4bit=True)
model = PeftModel.from_pretrained(base, "lokeshrao226/eco-logistics-qwen-grpo-v2")
tokenizer = AutoTokenizer.from_pretrained("lokeshrao226/eco-logistics-qwen-grpo-v2")For full inference + parser: see inference.py in the GitHub repo.
Live Environment
The supply-chain simulator this LoRA was trained against runs as a Hugging Face Space:
https://huggingface.co/spaces/lokeshrao226/eco-logistics
Hit /reset_v2, /submit_plan, /run_chunk, /grader to evaluate.
Limitations
- Constrained decoding is part of the policy. Direct deployment of this LoRA without the post-hoc parser will hit Air-shipment carbon overshoots and degrade by ~50% on the hard task.
- GRPO training was volatile (two collapse-recover cycles in 40 steps). We report held-out eval as ground truth.
- Single-agent only. The v2 environment supports 3-role multi-agent coordination, but this LoRA was trained single-agent (
role=solo). - No-op exploit on net_zero_profit grader. A degenerate "ship nothing" policy posts grader 0.465 on this task because the grader rewards profit-under-budget and idle inventory still partially fulfills demand. Heuristic and v9 both beat v8 on this grader; all are below the no-op exploit due to grader geometry, not policy quality. We flag this honestly.
Citation
@misc{eco-logistics-qwen-grpo-v2,
author = {J Lokesh Rao (Crystal Blue)},
title = {Eco-Logistics: Multi-City Supply Chain Optimizer (v9 LoRA)},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/lokeshrao226/eco-logistics-qwen-grpo-v2}
}