CoolFace
Apppublic

SabaPivot/repro-on-the-power-of-approximate-reward-models-for-inference-time-scaling-sequential-monte-carl

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
App README

Reproduction: On the Power of (Approximate) Reward Models for Inference-Time Scaling: Sequential Monte Carlo and Beyond

An open experiment logbook, published with Trackio.

Independent finite-state verification v2

The canonical reproduction now includes a second six-gate audit that replaces formula-only plots with actual hidden-prefix searches, exact finite-particle output laws, exhaustive small-policy enumeration, and a literal Metropolis–Hastings transition matrix.

  • —Re-run: python upgrade_smc_v2.py
  • —Results: outputs_v2/smc_v2_results.json
  • —Integrity manifest: outputs_v2/SHA256SUMS.json
  • —Result SHA-256: ac5e2396b6a75e0a148c58de9c106b1644211f1d4d033502328a105d235e8f37

All six evidence gates pass. A higher-scoring public reproduction was consulted to identify missing evidence categories; the implementation and all generated results here were independently produced.