openevo-recovery/openevo-qwen25-7b-webshop-sd-lora
OpenEVO Asset — openevo-qwen25-7b-webshop-sd-lora
<!-- openevo-cross-platform-identity:begin -->
中文说明(默认)
这是 OpenEVO 跨平台实验资产的一部分。本段只补充统一身份与导航信息,不修改仓库已有模型、数据、checkpoint、trajectory 或历史科学语义。
- Provider ID:
miyuki17/openevo-qwen25-7b-webshop-sd-lora - Role:
PUBLIC_FROZEN_SD_LORA - Classification:
historical-frozen-policy-artifact - Canonical:
true - Visibility:
public - Cross-platform binding:
registry-resolved - Retention:
KEEP_STABLE_ID - Scientific authority: pinned Git design/preregistration → run manifest/receipt/reconciliation;Hugging Face 是 durable artifact/provider layer。
- Governance:
mykcs/openevo-experiment@bb57c88eb5de3036dee3ea9b095bc550dd6882c8 - Language standard:
mykcs/openevo-experiment@a1994980f16a922374f25442a599031825b2cf03
命名和语言约定:repo、title、canonical ID 使用英文;README 默认中文。本“中文说明(默认)”是默认阅读入口,下方已有英文说明作为 English companion / historical detail 保留。若本段与 immutable publication receipt 或精确 remote revision/hash 冲突,以后者为准。
English: This block adds cross-platform identity metadata only. Existing artifact bytes and scientific claims are unchanged. <!-- openevo-cross-platform-identity:end -->
中文模型卡
导航 / Navigation:OpenEvo WebShop 公开科研产物
这是基于 `Qwen/Qwen2.5-7B-Instruct` 的 冻结 SD-LoRA 适配器(rank 32),由 OpenEvo H1.38B method-control 战役 (20260820-0129-h138b-method-control)产出,并在 SEED 官方 held-out WebShop 对比 (20260825-0624-webshop-seed-official-heldout-comparison)中作为 OPEN_EVO_SD 臂被评测。
这是一个仅推理的评测产物:适配器在此冻结、原样发布,供独立 Agent 复现/审阅 held-out WebShop 测量。没有再训练、没有调参、没有扫参。
可复现身份(使用前必校验)
把本适配器当作冻结政策使用前,务必确认 adapter_model.safetensors 的 SHA-256 等于 47e63d41… —— 这正是评测所 pin 的那个产物。
输出格式注意(WebShop 必看)
本政策在 <think> 块之后,用 `[action]…[/action]`(偶见残缺开标签 [action>)包裹动作。 SEED 官方 webshop_projection 只认 `<action>…</action>`,找不到时会退化成取末尾 20 字符的 碎片,导致每个 episode 都得 0。评测时请使用同时兼容两种 wrapper 的 projection(见评测仓库)。
评测结果(SEED 官方 held-out,env.seed=0,goal_idx 0–499,128 任务面板)
- SEED-strict primary 是测量无效:SEED 解析器读不了本政策的
[action]包裹,两臂 虽吐出了合法search[...]/click[...]指令却全记 0 —— 这不是政策真实水平。 - 有效的本地测量是 OpenEvo-native diagnostic:本适配器把连续 task score 约翻倍 (14.6 → 28.3),但不提升 exact success —— 更多“部分进度”,而非更多“完成购买”。
- SEED 论文参考值(Table 1,Qwen2.5-7B-Instruct):89.7 分 / 78.1% 成功率。这是论文 reported 数字;SEED 训练/checkpoint 未在本地复现,128 面板是 SEED-compatible,并非 论文确切分母。
完整证据(全部 512 条 episode、reconciliation、analysis)在评测仓库 mykcs/openevo-experiment 的 docs/evidence/seed-official-heldout-comparison-v1/ (见 RESULTS.md)。
训练摘要
- 算法族:SD-LoRA —— continual SFT(
causal_lm_continual_sft_v4),bounded trajectory replay + 冻结全局单位 Frobenius 方向。 - 有效 rank:32;目标模块:
q_proj、v_proj;训练峰值显存约 16.4 GB。
边界与限制
- 单一冻结适配器,非模型扫参;仅推理评测。
- 未复现 SEED 89.7;不做完全对齐或因果性声明。
- held-out 面板 exact success 较低(≤2.3%);适配器提升的是部分进度,不是完成购买。
引用
使用本产物请引用 OpenEvo 实验仓库(mykcs/openevo-experiment)与 Qwen2.5 基座模型。
OpenEvo SD-LoRA — Qwen2.5-7B-Instruct (WebShop continual SFT)
中文说明见下方「中文模型卡」一节。This card is bilingual; the Chinese section below carries the same content.
A frozen SD-LoRA adapter (rank 32) on `Qwen/Qwen2.5-7B-Instruct`, produced by the OpenEvo H1.38B method-control campaign (20260820-0129-h138b-method-control) and evaluated as the OPEN_EVO_SD arm in the SEED official held-out WebShop comparison (20260825-0624-webshop-seed-official-heldout-comparison).
This is an inference-only evaluation artifact: the adapter was trained earlier and is published here frozen and unmodified so that independent agents can reproduce / audit the held-out WebShop measurement. No further training, tuning, or sweeping was done.
Reproducibility identity (verify before use)
Always confirm the adapter_model.safetensors SHA-256 matches 47e63d41… before using this as the frozen policy — this is the exact artifact the evaluation pinned.
Files
Only adapter_model.safetensors + adapter_config.json are needed for inference; the openevo_sd_lora_* files document how the adapter was produced.
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct", device_map="auto")
model = PeftModel.from_pretrained(base, "miyuki17/openevo-qwen25-7b-webshop-sd-lora")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")Output-format note (important for WebShop): this policy emits its chosen action wrapped in [action]…[/action] (sometimes a malformed [action> open tag) after a <think> block. If you evaluate it with SEED's released webshop_projection, that parser only recognises <action>…</action> and will silently fall back to a 20-character tail fragment, scoring every episode 0. Use a projection that accepts both wrappers (see the evaluation repo).
Evaluation (SEED official held-out WebShop, env.seed=0, goal_idx 0–499, 128-task panel)
- The SEED-strict primary result is measurement-invalid: SEED's parser cannot read this policy's
[action]wrapper (see Output-format note), so both arms score 0 despite emitting validsearch[...]/click[...]commands. It is not a policy result. - The valid local measurement is the OpenEvo-native diagnostic: the adapter roughly doubles continuous task score (14.6 → 28.3) but does not improve exact success — more partial progress, not more completed purchases.
- SEED paper-reported reference (Table 1, Qwen2.5-7B-Instruct): 89.7 score / 78.1% success. This is a paper number; SEED training/checkpoint was not locally reproduced, and the 128-task panel is SEED-compatible, not the exact paper denominator.
Full evidence (all 512 episode records, reconciliation, analysis) lives in the evaluation repository: mykcs/openevo-experiment → docs/evidence/seed-official-heldout-comparison-v1/ (see RESULTS.md).
Training summary
- Algorithm family: SD-LoRA — continual SFT (
causal_lm_continual_sft_v4) with bounded trajectory replay and a frozen global unit-Frobenius direction. - Effective rank: 32; target modules:
q_proj,v_proj. - Peak GPU memory during training: ~16.4 GB.
- Source campaign: H1.38B method-control (
20260820-0129-h138b-method-control).
Limitations & claim boundary
- A single frozen adapter, not a model sweep; evaluated inference-only.
- Does not reproduce SEED's 89.7; no exact apples-to-apples or causal claim is made.
- Exact success on the held-out panel is low (≤2.3%); the adapter improves partial task progress, not completed purchases.
Citation
If you use this artifact, cite the OpenEvo experiment repository (mykcs/openevo-experiment) and the Qwen2.5 base model.
