datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arc-whestbench-public-2026
Organized by:
Alignment Research Center (ARC),
AIcrowd
WhestBench 2026: ARC White-Box Estimation Challenge
WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs.
This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/arc-whestbench-public-2026.whestbench-smoke-mlp
Organized by:
Alignment Research Center (ARC),
AIcrowd
WhestBench 2026: ARC White-Box Estimation Challenge
WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs.
This is the train dataset for… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/whestbench-smoke-mlp.whestbench-ci-fixture
Organized by:
Alignment Research Center (ARC),
AIcrowd
WhestBench 2026: ARC White-Box Estimation Challenge
/tree/main">
WhestBench is a benchmark for white-box activation estimation: given the weights of a small ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs.
This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/whestbench-ci-fixture.whestbench-relu-mlp-moments-10k
WhestBench Random ReLU MLPs — Monte-Carlo Activation Cumulants (10k)
A dataset of 10,500 random ReLU MLPs (10,000 train + 500 held-out test) together with
Monte-Carlo–estimated per-layer activation cumulants (mean, variance, skewness, kurtosis) for
every layer, for both the pre-activation and post-ReLU signals. Built for the WhestBench
Estimation Challenge 2026 and for research on analytic moment / uncertainty propagation
through deep networks.
The generative process… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/whestbench-relu-mlp-moments-10k.arc-whestbench-convergence-2026-traj-b
WhestBench convergence study — v1-warmup-100mlps
Per-neuron cumulative mean activations of the first 100 MLPs of
aicrowd/arc-whestbench-public-2026
captured at 860 log-spaced sample budgets from
N=1 to N=1,000,000,000.
This is a research sidecar to v1-warmup. It demonstrates how Monte Carlo mean
estimates of the activations converge as the sample budget grows.
This is trajectory B — an independent peer of
aicrowd/arc-whestbench-convergence-2026
(trajectory A). Same 100 MLPs… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/arc-whestbench-convergence-2026-traj-b.
