mrzhao13/qwen3.5-2b-shopsimulator-sft-512-1ep
Qwen3.5-2B ShopSimulator SFT-512 1 Epoch
This is a Hugging Face export of Qwen/Qwen3.5-2B after one epoch of supervised fine-tuning on successful ShopSimulator teacher trajectories. Training used the Pi agent harness and the modified Slime integration.
The fine-tuning data is text-only. The checkpoint retains the original multimodal architecture, but visual capabilities were not trained or evaluated in this experiment.
Training
The training dataset repository is currently private while the redistribution status of the upstream ShopSimulator task content is being clarified.
Evaluation
The following results use the same modified ShopSimulator service, deterministic price generation, fixed official_test_200 task slice, and one rollout per task. These are k=1 point estimates, not uncertainty estimates.
Usage
Use this repository as a drop-in replacement for Qwen/Qwen3.5-2B with a recent Transformers version that supports Qwen3.5. Refer to the official Qwen3.5-2B model card for loading and inference examples, replacing the model ID with this repository.
For ShopSimulator agent evaluation, use the Pi extension, prompt template and environment adapter from the companion pi-slime-shopsimulator source repository when it is published.
Limitations
- Fine-tuned only on Chinese ShopSimulator trajectories.
- Evaluated only in the patched text-based ShopSimulator environment.
- The deterministic pricing patch changes scores relative to the unpatched upstream environment.
- The official evaluation contains 200 tasks with one stochastic rollout each.
- Visual behavior, general instruction following and deployment safety were not evaluated.
- Teacher-generated content may contain errors.
License and attribution
The base model and these derivative weights are distributed under Apache-2.0. The original Qwen license is included in LICENSE. This model card identifies the checkpoint as modified and does not grant rights to third-party ShopSimulator task or environment content.
