linalg
Datasets
All datasets matching “linalg”LinAlgBench
LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning
A diagnostic benchmark evaluating 10 frontier LLMs on linear algebra computation across a strict dimensional gradient (3×3, 4×4, 5×5 matrices), with automated three-stage forensic error classification revealing that LLM mathematical failure is structurally constrained by algorithm family and matrix dimension rather than random.
Overview
LinAlg-Bench is a diagnostic dataset… See the full description on the dataset page: https://huggingface.co/datasets/linalgbench2026/LinAlgBench.lin-alg-kernels-corelinalg-bench
LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning
A diagnostic benchmark evaluating 10 frontier LLMs on linear algebra computation across a strict dimensional gradient (3×3, 4×4, 5×5 matrices), with automated three-stage forensic error classification revealing that LLM mathematical failure is structurally constrained by algorithm family and matrix dimension rather than random.
Overview
LinAlg-Bench is a diagnostic… See the full description on the dataset page: https://huggingface.co/datasets/LinAlgBench/linalg-bench.linalg-bench-llm
LinAlg-Bench: Where LLMs Stop Computing and Start Hallucinating
Ten frontier LLMs drop from near-perfect to near-zero on 5×5 eigenvalue problems. Complete computational collapse is dimension-gated: rare at 3×3, dominant at 4×4 and 5×5. Failures dissociate cleanly by task — eigenvalues fail by constraint-aware fabrication (invented eigenvalues that still match the matrix trace), determinants by sign-accumulation drift. Nearly a third of irrational-spectrum eigenvalue failures are… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-llm.LinAlg-Model-Outputs
LinAlg-Bench: Linear Algebra Problem Solving Benchmark for LLMs (Reviewer Preview)
Reviewer Preview — Limited Data Release. This repository is a limited-scope preview prepared for NeurIPS 2026 reviewer access, not the final camera-ready dataset. It contains the benchmark problems, raw model responses across three independent evaluation runs, and multi-run accuracy summaries only.
A diagnostic benchmark evaluating 10 frontier LLMs on linear algebra computation across a strict… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/LinAlg-Model-Outputs.linalgzero-distilled-failures
Dataset Card for linalgzero-distilled-failures
Information about how this dataset is used is available in the linalg-zero repository.
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config… See the full description on the dataset page: https://huggingface.co/datasets/rfvasile/linalgzero-distilled-failures.
