CoolFace
Datasetpublic

reasoning-degeneration-dev/t1-strategy-musr-baseline-qwen3-80b-thinking-musr-base

t1-strategy-musr-baseline-qwen3-80b-thinking-musr-base Strategy compliance evaluation on MuSR murder mysteries — baseline variant. Model received standard MuSR cot+ prompt (baseline control). Judge still scores against criterion-first rubric. Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric. Results Metric Value pass@1 0.9000 Strategy compliance (mean) 2.00 Strategy compliance (min) 2 Strategy compliance (max) 2… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-baseline-qwen3-80b-thinking-musr-base.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
0likes8downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
reasoning-degeneration-dev/t1-strategy-musr-baseline-qwen3-80b-thinking-musr-base · CoolFace