CoolFace
Datasetpublic

violetxi/Qwen3.5-9B-EquationalTheories-Proof-nl-eval

Qwen3.5-9B base: Equational Theories prose-proof evaluation Accuracy: 42.73% (282/660 correct). Single-turn natural-language proof correctness, graded by GPT-5.6-sol. These are model-judged results; generated proofs were not checked by the Lean kernel. Model Accuracy Correct / total Change vs base Base (this dataset) 42.73% 282/660 +0.00 pp 1M 38.64% 255/660 -4.09 pp 3M 44.85% 296/660 +2.12 pp 10M 43.64% 288/660 +0.91 pp The default dataset viewer now shows… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/Qwen3.5-9B-EquationalTheories-Proof-nl-eval.

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes64downloads

violetxi/Qwen3.5-9B-EquationalTheories-Proof-nl-eval · main · files are served by the source, never re-hosted here