ToniDO/TeXtract_dataset
TeXtract_dataset (WebDataset Format) This repository contains approximately 3.2 million pairs of mathematical expression images and their corresponding LaTeX source code, packaged in WebDataset format for large-scale training. The dataset is based on and derived from the original hoang-quoc-trung/fusion-image-to-latex-datasets, transformed for more efficient access. 📂 Dataset Structure Each WebDataset shard (.tar) contains multiple samples. Each sample groups… See the full description on the dataset page: https://huggingface.co/datasets/ToniDO/TeXtract_dataset.
This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.
