signaldepth/e15-context-budget
SignalDepth E15 Context Budget This is a small prompt-sensitivity benchmark slice for separating two explanations that often get conflated: the prompt is too short the task contract is underspecified The narrow result: on this deterministic Python code-task suite, making sparse prompts longer did not help. Making the task contract explicit did. Key Result Condition Average pass rate Read short_sparse 0.25 short and underspecified long_sparse 0.25… See the full description on the dataset page: https://huggingface.co/datasets/signaldepth/e15-context-budget.
010
