psikosen/canopy-258m-r3-v8
Canopy-258M-R3 v8
Experimental BF16 comparison-specialized checkpoint, initialized from v7. Selected arm: independent_orders. Published at the user's request; broader arithmetic, copying, JSON, conversation and executed-code retention checks remain pending.
Run
hf download psikosen/canopy-258m-r3-v8 --local-dir canopy-v8
cd canopy-v8
pip install -r requirements.txt
python run.py 'Which number is larger: 17 or 38? Answer with only the number.'Uses bundled native PyTorch inference, not Transformers AutoModel remote-code loading. Evaluated with BF16 weights, full attention, greedy decoding, a 32-token ceiling, repetition penalty 1 and stochastic recurrence disabled. These are not packed 1.58-bit weights.
Controlled comparison experiment
The selected run improved this narrow test by 20.3 percentage points. Both runs used 120 SFT steps, answer-averaged loss, the same order-balanced training pool and chat/code replay. Selection preceded testing; independent_orders won the validation tie by fixed arm order. Pairing both orders within each batch did not demonstrate additional benefit.
There are only 32 independent test pairs, each asked in both orders. The test wording was absent from training; validation used that wording on different pairs. One training seed. Historical pretraining contamination was not exhaustively excluded. These scores are not general benchmark accuracy, browser performance, or code-execution results.
Retention-anchor NLL changed from 3.18294 to 3.20003; this is a loss proxy on 24 historical examples, not proof that broader capabilities were preserved. Earlier v7 arithmetic and copying weaknesses should be assumed unresolved until tested. Raw results and the pre-publication experiment notes are in evaluation/. The training script there is an archival source snapshot and depends on the original project's v7 training helpers.
Weight SHA256: a18a9700ad0048f51eebb2988912019ea9206bc5540347e7b6c68b01732a47e8
