datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arXiv_latex
TeX data from arXiv
Using https://github.com/KyuDan1/TeX2Image code.
We have categories Math, Physics, Statistics, ComputerScience.
Domain
Size
Mathematics
4.22M
Computer Science
2.76M
Statistics
0.89M
Physics
0.78M
Total (unique)
7.17M
fusion-image-to-latex-datasets
Collects and builds the largest dataset to date from online sources, creating a robust and generalizable dataset. This dataset includes approximately 3.4 million image-text pairs, including both handwritten mathematical expressions (200,330 examples) and printed mathematical expressions (3,237,250 examples). Due to the large dataset and the fact that the same mathematical formula can be represented in different LaTeX string formats in an image, it is easy to cause polymorphic ambiguity. To… See the full description on the dataset page: https://huggingface.co/datasets/hoang-quoc-trung/fusion-image-to-latex-datasets.LaTeXLatex_Nemeth
