CoolFace
Datasetpublic

Anon987281293/ProofRank-outputs

ProofRank: Evaluation Outputs Companion artifact to the NeurIPS 2026 submission "Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness". This dataset contains the complete LLM-judge outputs behind every number reported in the paper, so that all results can be recomputed without re-querying any model. It covers the ten evaluated models (GPT-5.4, Gemini-3.1-Pro, Gemini-3-Flash, GLM-5, DeepSeek-v3.2, Kimi-K2.5-Think, StepFun-3.5-Flash, Qwen3.5-397B… See the full description on the dataset page: https://huggingface.co/datasets/Anon987281293/ProofRank-outputs.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes258downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Anon987281293/ProofRank-outputs · CoolFace