GroupieSteven/fullstackarena-gpt56sol-rideapp-selective-v812-20260912
RideApp GPT-5.6-sol selective difficulty calibration This gated evidence release contains the final observation selected for one representative of each of RideApp's 30 task templates. It is not a fresh 30-task run. Following the registered PI-directed protocol, the 18 templates that passed the September 5 baseline were rerun after difficulty refinement; the 12 baseline failures were carried forward without another model call. Three policy-sensitive observations and two… See the full description on the dataset page: https://huggingface.co/datasets/GroupieSteven/fullstackarena-gpt56sol-rideapp-selective-v812-20260912.
018
No commit history came back for main. The revision may not exist, or the source declined the request.
