evalstate/cifar10-selftrain-best-iterate
011
CIFAR-10 Self-Training Best (iterate t=5)
Small CNN (MediumCNN(4 conv blocks 64/128/256/512 + GAP, ~4.69M params); not ResNet-50) trained as iterate t=5 in a self-distillation trajectory reproducing the denoising-vs-forgetting trade-off from Wu, Yang & Sun, Why Self-Training Helps and Hurts (arXiv:2602.14029).
This iterate
- test accuracy: 55.76% (error 44.24%)
- mean confidence: 0.889
- agreement with previous iterate's pseudo-labels: 56.02%
- label source: hard pseudo-labels from iterate 4
Trajectory config (run seed=0)
- Dataset: CIFAR-10 (uoft-cs/cifar10), disjoint n=5000 subset per iterate
- Label noise: symmetric, eta=0.4 (teacher t=0 only) -> injects initial stochastic error
- Students t>=1: trained from scratch on HARD pseudo-labels from the previous iterate
- Optimizer: SGD lr=0.1 mom=0.9 wd=0.0005 nesterov + cosine, 50 epochs, batch 256
- Augmentation: RandomCrop(pad4,reflect)+HFlip+Normalize (GPU)
- Best iterate of the trajectory: t*=5 (acc 55.76%)
This is a budget-conscious qualitative reproduction (NOT paper-scale ResNet-50). The full trajectory exhibits a U-shaped test-error curve (denoise then forget). Full per-iterate trajectory CSV: see evalstate/synthetic-selftrain-denoising-forgetting.
