CoolFace
Modelpublic

lokeshrao226/eco-logistics-qwen-grpo-v2

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes8downloads
Model Card

Eco-Logistics Qwen-2.5-1.5B GRPO (v9)

GRPO-trained LoRA adapter for Qwen-2.5-1.5B-Instruct that plans 25-step supply-chain trajectories under a strict carbon budget.

Submitted to OpenEnv Hackathon Round 2 (Crystal Blue) — World Modeling for Professional Tasks.

Headline Numbers

Held-out eval on 10 seeds × 3 tasks = 30 episodes, greedy decoding, with constrained parser:

Taskv9 (this)v8 (prior submission)
Restock Only (easy)0.9990.13
Inventory Balanced (medium)0.2690.19
Net-Zero Profit (hard)0.3380.273
Valid action rate100%20%

Architecture

  • —Base: Qwen/Qwen2.5-1.5B-Instruct, 4-bit quantized
  • —Adapter: LoRA r=16, target modules [q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj]
  • —Trainable parameters: 18.5M (1.18% of base)

Training

  • —Method: GRPO via TRL 0.24, Unsloth wrapping, single T4 GPU
  • —Reward: grader-only (avoids the SFT-then-GRPO collapse seen in v8 experiments)
  • —Steps: 40 GRPO steps, ~58 min on T4
  • —Hyperparameters: 4 generations/prompt, batch 4, grad_accum 4, lr 1.5e-6
  • —Prompts: 80 unique initial states (seeds 0–119)
  • —Task mix during training: 60% netzeroprofit, 40% inventory_balanced

Inference (CRITICAL: use constrained parser)

The model is paired with a constrained parser that projects model output to feasible actions:

python
def parse_plan(completion, target_length=25):
    # ... extract JSON array ...
    cleaned.append({
        "ship_amount": min(2.0, max(0.0, float(item.get("ship_amount", 0.0)))),
        "origin_city": _coerce_city(item.get("origin_city", ""), "Seattle"),
        "destination_city": _coerce_city(item.get("destination_city", ""), "Chicago"),
        "speed_mode": "Rail",   # FORCED — Air carbon kills the budget
    })

The cap ship≤2 is derived from carbon-budget arithmetic: 25 steps × 2 ship × 2 rail-carbon ≤ 100 budget.

Without the parser, raw model output achieves grader 0.144 on net_zero_profit. With the parser, 0.338. Treat the LoRA + parser as a single policy.

Quick Use

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct", load_in_4bit=True)
model = PeftModel.from_pretrained(base, "lokeshrao226/eco-logistics-qwen-grpo-v2")
tokenizer = AutoTokenizer.from_pretrained("lokeshrao226/eco-logistics-qwen-grpo-v2")

For full inference + parser: see inference.py in the GitHub repo.

Live Environment

The supply-chain simulator this LoRA was trained against runs as a Hugging Face Space:

https://huggingface.co/spaces/lokeshrao226/eco-logistics

Hit /reset_v2, /submit_plan, /run_chunk, /grader to evaluate.

Limitations

  • —Constrained decoding is part of the policy. Direct deployment of this LoRA without the post-hoc parser will hit Air-shipment carbon overshoots and degrade by ~50% on the hard task.
  • —GRPO training was volatile (two collapse-recover cycles in 40 steps). We report held-out eval as ground truth.
  • —Single-agent only. The v2 environment supports 3-role multi-agent coordination, but this LoRA was trained single-agent (role=solo).
  • —No-op exploit on net_zero_profit grader. A degenerate "ship nothing" policy posts grader 0.465 on this task because the grader rewards profit-under-budget and idle inventory still partially fulfills demand. Heuristic and v9 both beat v8 on this grader; all are below the no-op exploit due to grader geometry, not policy quality. We flag this honestly.

Citation

bibtex
@misc{eco-logistics-qwen-grpo-v2,
  author = {J Lokesh Rao (Crystal Blue)},
  title  = {Eco-Logistics: Multi-City Supply Chain Optimizer (v9 LoRA)},
  year   = {2026},
  publisher = {Hugging Face},
  url    = {https://huggingface.co/lokeshrao226/eco-logistics-qwen-grpo-v2}
}