fimu-docproc-research/CIVQA-TesseractOCR
CIVQA TesseractOCR Dataset The Czech Invoice Visual Question Answering dataset was created with Tesseract OCR, and it is suitable for adding labels for the chosen model. The encoded dataset for LayoutLM model can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR-LayoutLM All invoices used in this dataset were obtained from public sources. Over these invoices, we were focusing on 15 different entities, which are crucial for… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR.
183
