CoolFace
Datasetpublic

bevaya/TABMEpp

Dataset Card for TABME++ The TABME dataset is a synthetic collection of business document folders generated from the Truth Tobacco Industry Documents archive, with preprocessing and OCR results included, designed to simulate real-world digitization tasks. TABME++ extends TABME by enriching it with commercial-quality OCR (Microsoft OCR). Dataset Details Dataset Description The TABME dataset is a synthetic collection created to simulate the… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/TABMEpp.

sourceHugging Facemitupdated 2y agoView on Hugging Face
5likes350downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
bevaya/TABMEpp · CoolFace