datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linalg-bench-llm
LinAlg-Bench: Where LLMs Stop Computing and Start Hallucinating
Ten frontier LLMs drop from near-perfect to near-zero on 5×5 eigenvalue problems. Complete computational collapse is dimension-gated: rare at 3×3, dominant at 4×4 and 5×5. Failures dissociate cleanly by task — eigenvalues fail by constraint-aware fabrication (invented eigenvalues that still match the matrix trace), determinants by sign-accumulation drift. Nearly a third of irrational-spectrum eigenvalue failures are… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-llm.Linalg-Spec-30
Linalg-Spec-30
Accepted to NeurIPS 2026 (Evaluations & Datasets Track)Paper (arXiv) · Code (GitHub) · All six datasets (collection)
Hand-authored NL→MLIR pairs for linalg named ops under memref semantics (n=30).
This dataset is one of six NL→MLIR benchmarks released with the NeurIPS 2026 Evaluations & Datasets Track paper Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR (arXiv:2607.18254). The full suite… See the full description on the dataset page: https://huggingface.co/datasets/plawanrath/Linalg-Spec-30.
