rwlinno/topoprm-ckpts
TopoPRM Checkpoints
Model artifacts for Rewarding the Graph Behind the Chain: Topology-Aware Process Supervision for RL and Reasoning Distillation.
Latest completed Full release
The September 26, 2026 release contains the completed Full run with seed 123. Stage III used 512 revision attempts, accepted 335, and completed 64 optimizer updates. Its model card records the training settings and checkpoint scope.
This checkpoint is a separate replication run from the seed-42 supplementary mechanism experiments and earlier main-result checkpoints. Benchmark scores from those runs must not be attached to this checkpoint. Its release does not include an evaluated benchmark score table.
Load the merged model
Use Transformers 5.5.4 or a compatible version. No PEFT adapter is needed for this release.
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "rwlinno/topoprm-ckpts"
folder = "topoprm-qwen35-9b-full-tgd-seed123-20260926"
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=folder)
model = AutoModelForCausalLM.from_pretrained(
repo, subfolder=folder, dtype="auto", device_map="auto"
)The model contains the text backbone only. Mathematical answers still require independent verification.
Earlier checkpoints
Existing SFT, GRPO, SCAE, and OPD adapter directories remain available as historical artifacts. They represent different training stages and protocols. They are not interchangeable with the latest fully merged Full TGD release. The repository-root adapter also remains a legacy artifact. Select the explicit subdirectory above to load the new model.
License
The new Qwen3.5-9B-derived Full release is distributed under Apache-2.0, with its license included in the checkpoint directory. Earlier checkpoints retain the terms applicable to their respective base models and releases.
