CoolFace
Datasetpublicgated

ToniDO/TeXtract_dataset

TeXtract_dataset (WebDataset Format) This repository contains approximately 3.2 million pairs of mathematical expression images and their corresponding LaTeX source code, packaged in WebDataset format for large-scale training. The dataset is based on and derived from the original hoang-quoc-trung/fusion-image-to-latex-datasets, transformed for more efficient access. 📂 Dataset Structure Each WebDataset shard (.tar) contains multiple samples. Each sample groups… See the full description on the dataset page: https://huggingface.co/datasets/ToniDO/TeXtract_dataset.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes3downloads
filemath_dataset-000000.tar2.21 GBdownload
filemath_dataset-000001.tar2.21 GBdownload
filemath_dataset-000002.tar2.21 GBdownload
filemath_dataset-000003.tar2.19 GBdownload
filemath_dataset-000004.tar2.21 GBdownload
filemath_dataset-000005.tar2.22 GBdownload
filemath_dataset-000006.tar2.21 GBdownload
filemath_dataset-000007.tar2.20 GBdownload
filemath_dataset-000008.tar2.20 GBdownload
filemath_dataset-000009.tar2.21 GBdownload
filemath_dataset-000010.tar2.22 GBdownload
filemath_dataset-000011.tar2.21 GBdownload
filemath_dataset-000012.tar2.19 GBdownload
filemath_dataset-000013.tar2.19 GBdownload
filemath_dataset-000014.tar2.19 GBdownload
filemath_dataset-000015.tar2.21 GBdownload
filemath_dataset-000016.tar139.4 MBdownload

ToniDO/TeXtract_dataset · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.