CoolFace
Datasetpublic

violetxi/Qwen3.5-9B-eq70-3m-v2-EquationalTheories-Proof-nl-eval

Qwen3.5-9B 3m: Equational Theories prose-proof evaluation Accuracy: 44.85% (296/660 correct). Single-turn natural-language proof correctness, graded by GPT-5.6-sol. These are model-judged results; generated proofs were not checked by the Lean kernel. Model Accuracy Correct / total Change vs base Base 42.73% 282/660 +0.00 pp 1M 38.64% 255/660 -4.09 pp 3M (this dataset) 44.85% 296/660 +2.12 pp 10M 43.64% 288/660 +0.91 pp The default dataset viewer now shows… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/Qwen3.5-9B-eq70-3m-v2-EquationalTheories-Proof-nl-eval.

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes67downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
violetxi/Qwen3.5-9B-eq70-3m-v2-EquationalTheories-Proof-nl-eval · CoolFace