CoolFace
Datasetpublic

lukbl/LaTeX-OCR-dataset

LaTeX-OCR Dataset Summary This dataset was created to train LaTeX-OCR, a model for recognizing LaTeX code from images of mathematical formulas. Each sample consists of a synthetically rendered formula image and its corresponding LaTeX formula. The images were generated from scratch using xelatex with multiple fonts, offering more typographic variety than many other datasets that use a single font (typically Computer Modern). Data Sources Formulas… See the full description on the dataset page: https://huggingface.co/datasets/lukbl/LaTeX-OCR-dataset.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes67downloads
6 commits on main
090b5151y ago

Update README.md

lukbl
b8fd1dd1y ago

Upload README.md with huggingface_hub

lukbl
433ee311y ago

Upload data/validation-00000-of-00001-79453adb1121b347.parquet with huggingface_hub

lukbl
8298dc41y ago

Upload README.md with huggingface_hub

lukbl
98932e71y ago

Upload data/train-00000-of-00001-5cb175040388de1e.parquet with huggingface_hub

lukbl
62b672b1y ago

initial commit

lukbl