datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository.
IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set
FormulaBank-28K
FormulaBank-28K
FormulaBank-28K is a deterministic procedural-audio corpus for audio representation pre-training. It contains 28,000 mono clips at 16 kHz, each exactly 10.24 seconds long.
Configuration
Version: C125-I224-v1
Formula classes: 125
Renderings per class: 224
Total clips: 28,000
Audio format: lossless 24-bit FLAC
Source: frozen AudioPG-Atomic-H7-C224-R0-Clean FormulaBank manifest
Each formula class specifies an acoustic rendering rule. Each rendering… See the full description on the dataset page: https://huggingface.co/datasets/KevinJustin/FormulaBank-28K.canonical-formulas-v1
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/canonical-formulas-v1
The canonical SZL formula registry — 21 pure, typed, no-IO Python formulas, the
matching Lean 4 obligation theorems, and the Codex-Kernel governed-loop composer.
Contents
File
What
code/python/formulas.py
21 canonical formulas, each… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/canonical-formulas-v1.latex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/OleehyO/latex-formulas.doclaynet-pt-enriched-formulamath_formula_retrieval
Dataset Card for MFR (Mathematical Formula Retrieval)
This dataset consists of formula pairs, classified as either mathematical equivalent or not.
Dataset Details
Dataset Description
Mathematical dataset based on 71 famous mathematical identities. Each entry consists of two identities (in formula or textual form), together with a label, whether the two versions describe the same mathematical identity. The false pairs are not randomly chosen, but… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_formula_retrieval.named_math_formulas
Dataset Card for NMF (Named Mathematical Formulas)
This dataset might be used to train a language model based math-retrieval system with a NSP-like task.
See also see ddrg/math_formula_retrieval as a derived dataset which associates two formulas.
You can find more information in MAMUT: A Novel Framework for Modifying Mathematical Formulas for the Generation of Specialized Datasets for Language Model Training.
Named Math Formulas
Mathematical dataset based on 71… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/named_math_formulas.wikipedia-latex-formulas-319k
Wikipedia LaTeX Formulas 319k Dataset
A curated collection of mathematical formulas from English Wikipedia, providing LaTeX source code paired with high-quality rendered images for training OCR and formula recognition models.
Dataset Description
Dataset Summary
This dataset represents a complete extraction of ~319k LaTeX mathematical formulas from English Wikipedia articles (October 2025 snapshot), filtered by visual complexity (score > 8) and renderability… See the full description on the dataset page: https://huggingface.co/datasets/piushorn/wikipedia-latex-formulas-319k.XBRL_Formula_Calculationfxbench
FxBench
FxBench is a spreadsheet formula-completion benchmark. Each row contains
one target formula and the corresponding workbook snapshot needed to solve it.
Dataset contents
This release contains 503 examples.
Columns in the Hugging Face Parquet split:
Column
Type
Description
id
string
Public example ID in the form fxbench-{i}.
function
string
Primary Excel function for the target formula.
formula
string
Ground-truth Excel formula. Empty for ABSTAIN rows… See the full description on the dataset page: https://huggingface.co/datasets/FormulaComletion/fxbench.stl_formulae_variantsmath_formulas
Mathematical Formulas (MF)
Mathematical dataset containing formulas based on the AMPS Khan dataset and the ARQMath dataset V1.3. Based on the retrieved LaTeX formulas, more equivalent versions have been generated by applying randomized LaTeX printing with this SymPy fork using Math Mutator (MAMUT). The formulas are intended to be well applicable for MLM. For instance, a masking for a formula like (a+b)^2 = a^2 + 2ab + b^2 makes sense (e.g., (a+[MASK])^2 = a^2 + [MASK]ab + b[MASK]2… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_formulas.math-formulas
Math Formulas QA
Deterministic synthetic math QA dataset generated with seed 1337.
Properties
2,000,000 unique rows
1,800,000 train
100,000 validation
100,000 test
100,000 rows per Parquet shard
Every row is validated before it is written
kind alternates between problem and solution
Columns
question
answer
text
question_tex
answer_tex
family
difficulty
kind
validated
validator
formula_hash64
Families
Arithmetic, fractions… See the full description on the dataset page: https://huggingface.co/datasets/aplominski/math-formulas.named_math_formulas_ft
Named Math Formulas - Fine-Tuning Dataset
This dataset is a version of Named Math Formulas (NMF) dedicated to be used as fine-tuning dataset, e.g., by using the GitHub Repository aieng-lab/transformer-math-evaluation.
In contrast to the full version of NMF, this dataset contains much more meta data that can be used for fine-grained evaluations.
Dataset Details
Dataset Description
Mathematical dataset based on 71 famous mathematical identities. Each… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/named_math_formulas_ft.pubchem-smiles-molecular-formuladetails_formulae__Dorflan
Dataset Card for Evaluation run of formulae/Dorflan
Dataset Summary
Dataset automatically created during the evaluation run of model formulae/Dorflan on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_formulae__Dorflan.red-pill-drug-discovery-formulation
🔴 RED-PILL
Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language
The first open instruction-tuning dataset for drug discovery & formulation development.
Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions.
⚡ Quick Start
from datasets import load_dataset
# Load the full dataset
ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.FormulaCascade_Benchmark
FormulaCascade Benchmark
This repository hosts the anonymous review release of FormulaCascade, a benchmark for long-horizon spreadsheet formula-dependency and formula-family reasoning in real .xlsx workbooks.
Current Review Release
The current upload contains the eval-200 release version selected from v5 Challenge/Core/Broad release strata:
FormulaCascade_Benchmark_eval_200_Release_version.tar.gz: compressed eval-200 release-version split.… See the full description on the dataset page: https://huggingface.co/datasets/anonymity11/FormulaCascade_Benchmark.GameTheory-Formulator
🎯 GameTheory-Formulator
1,215 real-world scenarios mapped to formal game theory models with step-by-step formulation, solution, and interpretation.
📋 Overview
GameTheory-Formulator is the Phase 3 dataset in the Alogotron Game Theory pipeline. While GameTheory-Bench teaches models to solve formal game theory problems, this dataset teaches them to formulate real-world strategic scenarios as formal games — the critical missing link between natural language… See the full description on the dataset page: https://huggingface.co/datasets/Alogotron/GameTheory-Formulator.FormulaSpeech_datasets
FormulaSpeech Datasets
This repository provides the official datasets for Formula-Speech, a framework for improving scientific formula verbalization in large speech language models for accessible learning.
FormulaSpeech focuses on helping end-to-end large speech language models (LSLMs) accurately read scientific formulas in spoken form. The datasets are designed for speech-enabled AI tutors, especially in accessible learning scenarios where blind or low-vision learners rely on… See the full description on the dataset page: https://huggingface.co/datasets/Stephen-Lee/FormulaSpeech_datasets.formulacode-all
FormulaCode is a live benchmark for evaluating the holistic ability of LLM agents to optimize codebases. FormulaCode consists of two parts: a pipeline to construct performance optimization tasks, and an execution harness that connects a language model to our terminal sandbox.
This dataset contains 1215 enriched performance optimization tasks derived from real open-source Python projects, spanning 76 months of merged PRs.
The dataset is continuously updated… See the full description on the dataset page: https://huggingface.co/datasets/formulacode/formulacode-all.FormulaReasoning
FormulaReasoning
This is a Chinese-English bilingual question-answering dataset, which includes the following subsets:
formulareasoning
formulareasoning_enhancement
Each subset has the following split:
train.json: Training data
HoF_test.json: Homogeneous formulas testing data
HeF_test.json: Heterogeneous formulas testing data
Field Descriptions
Field
Type
Description
id
str
Each sample's unique identifier.
question
dict
Sample's question includes the… See the full description on the dataset page: https://huggingface.co/datasets/cat-overflow/FormulaReasoning.latex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/yayossd/latex-formulas.OleehyO-latex-formulasnamed_math_formulas
Dataset Card for NMF (Named Mathematical Formulas)
This dataset might be used to train a language model based math-retrieval system with a NSP-like task.
See also see ddrg/math_formula_retrieval as a derived dataset which associates two formulas.
You can find more information in MAMUT: A Novel Framework for Modifying Mathematical Formulas for the Generation of Specialized Datasets for Language Model Training.
Named Math Formulas
Mathematical dataset based on 71… See the full description on the dataset page: https://huggingface.co/datasets/shreyasmind/named_math_formulas.latex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Evanstarcraft2/latex-formulas.dom-formula-assignment-data
DOM Formula Assignment Dataset
Training and Testing Data for A Machine Learning and Benchmarking Approach for Molecular Formula Assignment of Ultra High-Resolution Mass Spectrometry Data from Complex Mixtures
Paper: Under review
Abstract
A machine learning approach to molecular formula assignment is crucial for unlocking the full potential of ultra-high resolution mass spectrometry (UHRMS) when analyzing complex mixtures. By combining data-driven models with… See the full description on the dataset page: https://huggingface.co/datasets/SaeedLab/dom-formula-assignment-data.formulae__mita-v1.1-7b-2-24-2025-details
Dataset Card for Evaluation run of formulae/mita-v1.1-7b-2-24-2025
Dataset automatically created during the evaluation run of model formulae/mita-v1.1-7b-2-24-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-v1.1-7b-2-24-2025-details.2024_formula1_championship_dataset
Formula 1 2024 Comprehensive LLM Fine-Tuning Dataset
🏁 Overview
This comprehensive dataset contains 507 high-quality training examples designed specifically for fine-tuning Large Language Models (LLMs) on Formula 1 2024 season data. This expanded version provides extensive coverage with multiple question variations and comprehensive F1 knowledge representation.
🏆 Training Example Categories
1. Race-Specific Questions (250+ examples)
Multiple… See the full description on the dataset page: https://huggingface.co/datasets/vibingshu/2024_formula1_championship_dataset.stl_formulaeThis dataset contains the data used in the paper "Bridging Logic and Learning: Decoding Temporal Logic Embeddings via Transformers" (Candussio et al. 2025).
More in detail:
train.csv is the training set for the random models. It contains formulae spanning from depth 2 to depth 23;
TODO: complete with test set.
