datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CIVQA_EasyOCR_Validation
CIVQA EasyOCR Validation Dataset
The CIVQA (Czech Invoice Visual Question Answering) dataset was created with EasyOCR. This dataset contains only the validation split. The train part of the dataset can be found on this URL: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_Train
The encoded validation dataset for the LayoutLM can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_Validation
All invoices used in this… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_Validation.CIVQA_EasyOCR_Train
CIVQA EasyOCR Train Dataset
The CIVQA (Czech Invoice Visual Question Answering) dataset was created with EasyOCR. This dataset contains only the train split. The validation part of the dataset can be found on this URL: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_Validation
The encoded train dataset for the LayoutLM can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_Train
All invoices used in this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_Train.qari-0.1-easyocr-eval-resultseasyocr_Qari-0.2-eval-tashkilCIVQA_EasyOCR_LayoutLM_Validation
CIVQA EasyOCR LayoutLM Validation Dataset
The CIVQA (Czech Invoice Visual Question Answering) dataset was created with EasyOCR, and it is encoded for LayoutLM models. This dataset contains only the validation split. The train part of the dataset can be found on this URL: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_TrainThe pre-encoded validation dataset can be found on this link:… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_Validation.CIVQA_EasyOCR_LayoutLM_Train
CIVQA EasyOCR LayoutLM Train Dataset
The CIVQA (Czech Invoice Visual Question Answering) dataset was created with EasyOCR, and it is encoded for LayoutLM models. This dataset contains only the train split. The validation part of the dataset can be found on this URL: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_ValidationThe pre-encoded train dataset can be found on this link:… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_Train.text_detections_easyocr_yolov10
