dvader13/forecastgen-artifacts
Forecast-Generalization: raw evaluation outputs across 38 reasoning models Complete generation-level outputs, per-seed scores and analysis artifacts from a study of how well benchmark performance forecasts generalization to held-out reasoning tasks. Most released evaluations report only aggregate accuracy. This release keeps the raw per-problem, per-seed generations, so item-level analyses can be redone without re-running any inference. What is here 38 models… See the full description on the dataset page: https://huggingface.co/datasets/dvader13/forecastgen-artifacts.
0193
