formula
Datasets
All datasets matching “formula”latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository.
IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set
FormulaBank-28K
FormulaBank-28K
FormulaBank-28K is a deterministic procedural-audio corpus for audio representation pre-training. It contains 28,000 mono clips at 16 kHz, each exactly 10.24 seconds long.
Configuration
Version: C125-I224-v1
Formula classes: 125
Renderings per class: 224
Total clips: 28,000
Audio format: lossless 24-bit FLAC
Source: frozen AudioPG-Atomic-H7-C224-R0-Clean FormulaBank manifest
Each formula class specifies an acoustic rendering rule. Each rendering… See the full description on the dataset page: https://huggingface.co/datasets/KevinJustin/FormulaBank-28K.canonical-formulas-v1
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/canonical-formulas-v1
The canonical SZL formula registry — 21 pure, typed, no-IO Python formulas, the
matching Lean 4 obligation theorems, and the Codex-Kernel governed-loop composer.
Contents
File
What
code/python/formulas.py
21 canonical formulas, each… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/canonical-formulas-v1.latex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/OleehyO/latex-formulas.doclaynet-pt-enriched-formulamath_formula_retrieval
Dataset Card for MFR (Mathematical Formula Retrieval)
This dataset consists of formula pairs, classified as either mathematical equivalent or not.
Dataset Details
Dataset Description
Mathematical dataset based on 71 famous mathematical identities. Each entry consists of two identities (in formula or textual form), together with a label, whether the two versions describe the same mathematical identity. The false pairs are not randomly chosen, but… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_formula_retrieval.
