test-time-scaling
fusion-pairwise-evals-test-time-scaling
Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N
Content
This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares CommandA against gemini-2.5-pro in 2 settings:
Test-time scaling with Fusion : 5 samples are generated from CommandA, then fused with CommandA into one completion and compared to a single completion from gemini-2.5-pro
Test-time scaling with BoN : 5 samples are generated from… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-test-time-scaling.testTimeScaling2parameters-researchs1-test-time-scaling-synth-public
s1-test-time-scaling-synth: Japanese and English Reinforcement Learning Dataset Derived from the s1 Simple Test-Time Scaling Dataset
This repository contains s1-test-time-scaling-synth, a reinforcement learning dataset in Japanese and English.This dataset is built upon the supervised fine-tuning dataset simplescaling/data_ablation_full59K (hereafter, the "original dataset"), originally developed in "s1: Simple test-time scaling" [Muennighoff+, EMNLP25].
The original dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/s1-test-time-scaling-synth-public.TestTimeScalingOlmo-1B-0724-hf-BON_50Q_profilingJ1-SFT-53k
J1-SFT-53k Dataset
Dataset Description
The J1-SFT-53k dataset is a supervised finetuning dataset for training LLM-as-a-Judge models, specifically curated for the J1-7B-SFT model described in the paper "J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge".
This dataset contains high-quality judgment examples enriched with reflective reasoning patterns.
Dataset Sources
This dataset is curated from multiple publicly available sources:
HelpSteer2
OffsetBias… See the full description on the dataset page: https://huggingface.co/datasets/test-time-scaling/J1-SFT-53k.
