whestbench
arc-whestbench-public-2026
Organized by:
Alignment Research Center (ARC),
AIcrowd
WhestBench 2026: ARC White-Box Estimation Challenge
WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs.
This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/arc-whestbench-public-2026.whestbench-smoke-mlp
Organized by:
Alignment Research Center (ARC),
AIcrowd
WhestBench 2026: ARC White-Box Estimation Challenge
WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs.
This is the train dataset for… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/whestbench-smoke-mlp.arc-whestbench-p2-full1000-N1e9
arc-whestbench Phase 2 — joint moments of the 1000 full-split MLPs at N = 10⁹
Monte-Carlo raw moments (marginal to 4th order, pairwise to 4th order) of the pre- and
post-activations of every one of the 1000 benchmark MLPs in
aicrowd/arc-whestbench-public-2026
(revision v2-phase2, config full), estimated from 10⁹ standard-normal input samples per
network. Built 2026-08-29 → 08-31 as the high-power replacement for the 100-network
mini moment sets; every quantity below is an… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/arc-whestbench-p2-full1000-N1e9.arc-whestbench-p2-higher-moments-2026
arc-whestbench-p2-higher-moments-2026
Higher-order joint activation moments for the Phase 2 (v2-phase2) benchmark MLPs
of the ARC WhiteBox Estimation Challenge 2026 — the 1024-wide, 16-layer analog of
arc-whestbench-higher-moments-2026.
Source MLPs: the 100 MLPs of the mini split of
aicrowd/arc-whestbench-public-2026
at revision v2-phase2 (width 1024, depth 16, He-init, no biases, fp32 weights,
input x ~ N(0, I), forward a_l = relu(a_{l-1} @ W_l)).
Moments are estimated from N =… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/arc-whestbench-p2-higher-moments-2026.whestbench-ci-fixture
Organized by:
Alignment Research Center (ARC),
AIcrowd
WhestBench 2026: ARC White-Box Estimation Challenge
/tree/main">
WhestBench is a benchmark for white-box activation estimation: given the weights of a small ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs.
This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/whestbench-ci-fixture.whestbench-relu-mlp-jointfeats-10k
whestbench-relu-mlp-jointfeats-10k
Companion to keenanpepper/whestbench-relu-mlp-moments-10k:
the joint (pairwise / third-cumulant-tensor) objects that the marginal-only
moments set does not contain, for the same 10,000 MLPs (bias-free ReLU,
width 256, depth 32, standard-normal input; seeds 0..9999, identical to
weights_train.npy in the moments set — index i ↔ seed i).
These are the features a deployable analytic estimator of the final-layer means
consumes: the marginal… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/whestbench-relu-mlp-jointfeats-10k.
