asingh15/tcs-qwen36-27b-direction-rollouts-pilot-50-full-trace
TCS Qwen3.6-27B Direction Beam Full-Trace Pilot Verified compact export for tcs_qwen36_27b_direction_beam_pilot50_full_logging_20260814. Source dataset: TCS train-00000-of-00001.parquet Problems: 50 Displayed trajectories: 200 Chunk probes: 1,600 Terminal answers and rubric grades: 6,400 Policy: Qwen/Qwen3.6-27B Judge: openai/gpt-oss-20b (low reasoning) At each displayed chunk, the policy proposes four directions plus a null continuation. A width-four stochastic beam reaches… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-rollouts-pilot-50-full-trace.
TCS Qwen3.6-27B Direction Beam Full-Trace Pilot
Verified compact export for tcs_qwen36_27b_direction_beam_pilot50_full_logging_20260814.
- Source dataset:
TCS train-00000-of-00001.parquet - Problems: 50
- Displayed trajectories: 200
- Chunk probes: 1,600
- Terminal answers and rubric grades: 6,400
- Policy:
Qwen/Qwen3.6-27B - Judge:
openai/gpt-oss-20b(low reasoning)
At each displayed chunk, the policy proposes four directions plus a null continuation. A width-four stochastic beam reaches four terminal solutions. Only terminal solutions are graded; the chunk value is their mean normalized rubric score. Browser bundles under web/ are split by problem, probe, and terminal so the explorer never downloads the full dataset.
