CoolFace
19 results

linalg

linalgbench2026 /LinAlgBench LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning A diagnostic benchmark evaluating 10 frontier LLMs on linear algebra computation across a strict dimensional gradient (3×3, 4×4, 5×5 matrices), with automated three-stage forensic error classification revealing that LLM mathematical failure is structurally constrained by algorithm family and matrix dimension rather than random. Overview LinAlg-Bench is a diagnostic dataset… See the full description on the dataset page: https://huggingface.co/datasets/linalgbench2026/LinAlgBench.1K<n<10K0 likes754 downloads5mo agoHugging FaceTokenBender /lin-alg-kernels-coretextn<1K0 likes639 downloads3mo agoHugging FaceLinAlgBench /linalg-bench LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning A diagnostic benchmark evaluating 10 frontier LLMs on linear algebra computation across a strict dimensional gradient (3×3, 4×4, 5×5 matrices), with automated three-stage forensic error classification revealing that LLM mathematical failure is structurally constrained by algorithm family and matrix dimension rather than random. Overview LinAlg-Bench is a diagnostic… See the full description on the dataset page: https://huggingface.co/datasets/LinAlgBench/linalg-bench.text10K<n<100K1 likes102 downloads4mo agoHugging Facemst-ai /linalg-bench-llm LinAlg-Bench: Where LLMs Stop Computing and Start Hallucinating Ten frontier LLMs drop from near-perfect to near-zero on 5×5 eigenvalue problems. Complete computational collapse is dimension-gated: rare at 3×3, dominant at 4×4 and 5×5. Failures dissociate cleanly by task — eigenvalues fail by constraint-aware fabrication (invented eigenvalues that still match the matrix trace), determinants by sign-accumulation drift. Nearly a third of irrational-spectrum eigenvalue failures are… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-llm.tabulartext-generation10K<n<100K0 likes94 downloads50m agoHugging Facemst-ai /LinAlg-Model-Outputs LinAlg-Bench: Linear Algebra Problem Solving Benchmark for LLMs (Reviewer Preview) Reviewer Preview — Limited Data Release. This repository is a limited-scope preview prepared for NeurIPS 2026 reviewer access, not the final camera-ready dataset. It contains the benchmark problems, raw model responses across three independent evaluation runs, and multi-run accuracy summaries only. A diagnostic benchmark evaluating 10 frontier LLMs on linear algebra computation across a strict… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/LinAlg-Model-Outputs.text10K<n<100K0 likes53 downloads2mo agoHugging Facerfvasile /linalgzero-distilled-failures Dataset Card for linalgzero-distilled-failures Information about how this dataset is used is available in the linalg-zero repository. This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config… See the full description on the dataset page: https://huggingface.co/datasets/rfvasile/linalgzero-distilled-failures.textn<1K0 likes44 downloads7mo agoHugging Face