violetxi/Qwen3.5-9B-eq70-3m-v2-EquationalTheories-Proof-nl-eval
Qwen3.5-9B 3m: Equational Theories prose-proof evaluation Accuracy: 44.85% (296/660 correct). Single-turn natural-language proof correctness, graded by GPT-5.6-sol. These are model-judged results; generated proofs were not checked by the Lean kernel. Model Accuracy Correct / total Change vs base Base 42.73% 282/660 +0.00 pp 1M 38.64% 255/660 -4.09 pp 3M (this dataset) 44.85% 296/660 +2.12 pp 10M 43.64% 288/660 +0.91 pp The default dataset viewer now shows… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/Qwen3.5-9B-eq70-3m-v2-EquationalTheories-Proof-nl-eval.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face