ToniDO/TeXtract_dataset
TeXtract_dataset (WebDataset Format) This repository contains approximately 3.2 million pairs of mathematical expression images and their corresponding LaTeX source code, packaged in WebDataset format for large-scale training. The dataset is based on and derived from the original hoang-quoc-trung/fusion-image-to-latex-datasets, transformed for more efficient access. 📂 Dataset Structure Each WebDataset shard (.tar) contains multiple samples. Each sample groups… See the full description on the dataset page: https://huggingface.co/datasets/ToniDO/TeXtract_dataset.
This repository is gated, so its file contents are only served once you have accepted the publisher's terms at Hugging Face. Open it at the source above.
