ToniDO/TeXtract_dataset
TeXtract_dataset (WebDataset Format) This repository contains approximately 3.2 million pairs of mathematical expression images and their corresponding LaTeX source code, packaged in WebDataset format for large-scale training. The dataset is based on and derived from the original hoang-quoc-trung/fusion-image-to-latex-datasets, transformed for more efficient access. 📂 Dataset Structure Each WebDataset shard (.tar) contains multiple samples. Each sample groups… See the full description on the dataset page: https://huggingface.co/datasets/ToniDO/TeXtract_dataset.
03
No card is published for this repository, or it could not be fetched from Hugging Face right now.
