CoolFace
Modelpublic

RampPublic/portal-qwen3-8b

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes50downloads
Model Card

PorTAL refit for Qwen3-8B

This is a native PorTAL artifact refitted onto Qwen/Qwen3-8B. Its 14-task latent table and canonical LoRA-generating core were learned jointly from Qwen3-1.7B and Qwen3-4B and frozen during refitting. Only a fresh Qwen3-8B alignment was trained. The artifact generates rank-8 LoRA factors for the query and value projections of every decoder layer.

Evaluation

One seed was evaluated on the complete 14-task validation suite using continuation log-probability divided by character length (acc_norm). Gold continuation token-mean NLL was tracked separately for checkpoint selection.

ModelMacro `acc_norm`
Frozen Qwen3-8B0.6681
PorTAL-adapted0.7767
Absolute lift+0.1086

These are research benchmark results for this exact artifact and evaluation recipe, not a general performance guarantee.

Refit recipe

  • —Base: Qwen/Qwen3-8B at b968826d9c46dd6066d109eabc6255188de91218
  • —Frozen source carrier: RampPublic/portal-qwen3-4b
  • —Dataset: RampPublic/portallib-tasks at ffc3c0e44f529bf64a5ae62ed5db090952db97ea
  • —Refit data: deterministic seeded sample of up to 1,000 examples per task from the complete training pool
  • —Optimization: 5 epochs, batch size 4, alignment LR 1e-3, linear decay with 10% warmup, seed 0
  • —Trainable parameters: target-base alignment only; task latents and canonical core remain frozen
  • —Checkpoint: maximum macro validation acc_norm, with lower gold NLL as the tie-breaker
  • —Architecture: q/v targets, rank 8, alpha 16, task latent 256, layer embedding 32, hidden 512, canonical width 1024

Usage

python
from portallib import PortalModel

portal = PortalModel.from_pretrained(
    "RampPublic/portal-qwen3-8b",
    revision="v0.2.0",
)
portal.export_peft("rte", "./portal-rte-qwen3-8b")

See the release recipe for the full task list, evaluation definition, and refitting procedure. The artifact is Apache-2.0; the benchmark dataset contains components under multiple upstream licenses documented on its dataset card.