CoolFace
Datasetpublic

JG1310/repro-chain-of-thought-gradient-descent-code-sol

Chain-of-Thought Gradient Descent — independent reproduction This directory contains a scaled, self-contained reproduction of the mechanisms behind ICML 2026 paper #443, Chain-of-Thought Gradient Descent. The paper does not provide code or complete training hyperparameters. This reproduction therefore tests its two central mechanisms directly: A frozen, one-layer decoder-style attention block is trained once and reused to emit local ReLU-network forward and backward blocks. It… See the full description on the dataset page: https://huggingface.co/datasets/JG1310/repro-chain-of-thought-gradient-descent-code-sol.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes37downloads
Dataset Card

Chain-of-Thought Gradient Descent — independent reproduction

This directory contains a scaled, self-contained reproduction of the mechanisms behind ICML 2026 paper #443, Chain-of-Thought Gradient Descent.

The paper does not provide code or complete training hyperparameters. This reproduction therefore tests its two central mechanisms directly:

  1. 1.A frozen, one-layer decoder-style attention block is trained once and reused to emit local ReLU-network forward and backward blocks. It receives exactly two routed state tokens per update and is evaluated by teacher forcing, autoregressive rollout, transfer to deeper networks, and recursive reuse across gradient-descent rounds.
  2. 2.A dynamic-mask benchmark compares processing two relevant weight matrices with processing all N matrices packed into every update. It reports both exact scalar-operation counts and measured device latency.

The experiment is intentionally scaled (d=4, small synthetic networks) and must be interpreted as a mechanism-level reproduction, not a full replication of the paper's unavailable training setup.

Run

bash
python scripts/run_reproduction.py --smoke --output-dir outputs/smoke
python scripts/run_reproduction.py --steps 3000 --batch-size 256 \
  --output-dir outputs/gpu
python scripts/make_figures.py --results outputs/gpu/results.json \
  --output-dir figures

All randomness is seeded. The main script records its configuration, package versions, device, wall-clock time, per-depth errors, multi-round rollout errors, dynamic-mask routing checks, exact cost counts, and timing measurements in one JSON file.

Reproduced result

The substantive run completed on one NVIDIA L4 in 912.1 seconds:

MetricN=3N=6 (unseen)N=9 (unseen)
Teacher-forced local component MSE0.0054640.0037530.003467
Round-1 free-running weight RMSE0.022510.0093060.008379
Round-10 free-running weight RMSE0.074330.063920.06631

The exact packed/masked processed-matrix ratio is N/2, hence Theta(N). Measured packed/masked L4 latency reached 4.42x at N=64 and 7.35x at N=128; this is lower than the exact operation-count ratio because fixed kernel overhead dominates at d=4.

  • —Successful GPU Job: <https://huggingface.co/jobs/JG1310/6a5cd3bfbee6ee1cf4ed1204>
  • —Immutable outputs: <https://huggingface.co/datasets/JG1310/repro-chain-of-thought-gradient-descent-runs-sol/tree/main/gpu-l4-seed443>
  • —Hub code mirror: <https://huggingface.co/datasets/JG1310/repro-chain-of-thought-gradient-descent-code-sol>
  • —Artifact collection: <https://huggingface.co/collections/JG1310/reproduction-chain-of-thought-gradient-descent-sol-6a5cd3845ba8f144af18312b>

Claim 1 and Claim 2 are therefore supported at toy scale/mechanism level. This does not establish the paper's arbitrary-depth theorem empirically, recreate its unavailable explicit attention weights, or reproduce its larger multi-seed Appendix E setup.