icdar
icdar23-entrydetector_plaintext_breaks_indents_left_diffflair-icdar-nldetr-resnet50_finetuned_icdar2019_finetuned_lstabledetv1hmbench-icdar-nl-hmbert_64k-bs8-wsFalse-e10-lr5e-05-poolingfirst-layers-1-crfFalse-2hmbench-icdar-nl-hmbert-bs4-wsFalse-e10-lr5e-05-poolingfirst-layers-1-crfFalse-2hmbench-icdar-nl-hmbert-bs8-wsFalse-e10-lr3e-05-poolingfirst-layers-1-crfFalse-5flair-icdar-nlicdar23-entrydetector_plaintext_breaks_indents_left_ref
Datasets
All datasets matching “icdar”ICDAR2019-SROIE
ICDAR2019's Scanned Receipts OCR and Information Extraction (SROIE)
The ICDAR2019 SROIE dataset was originally published by Huang et al. for the
15th International Conference on Document Analysis and Recognition (ICDAR2019)
Robust Reading Challenge on Scanned Receipts OCR and Information Extraction
(SROIE).
This work presents an extension of the original ICDAR2019 SROIE dataset, including 14
receipt annotations missing from the original Task 3 test dataset, in a format
integrated… See the full description on the dataset page: https://huggingface.co/datasets/jsdnrs/ICDAR2019-SROIE.ICDAR2015icdar2021-historical-document-dating
ICDAR 2021 Historical Document Classification — Task 2 (Dating)
13,810 manuscript page images labelled with the period in which they were produced.
Images come from e-codices, the virtual manuscript library
of Switzerland.
Split
Images
Date range
Median span
Dated to a single year
train
11,294
800–1899
45 years
1,409
test
2,516
800–1921
49 years
264
The label is an interval, not a year
Palaeographers date a manuscript to a range, and the width of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/icdar2021-historical-document-dating.icdar_disco
ICDAR_mini Dataset
A balanced mini subset of the ICDAR (International Conference on Document Analysis and Recognition) dataset with 50 samples per language. Includes actual document images and ground truth OCR text.
Dataset Details
Total Samples: 500
Total Images: 500
Languages: 10
Arabic (50 samples)
Bangla (50 samples)
Chinese (50 samples)
Hindi (50 samples)
Japanese (50 samples)
Korean (50 samples)
Latin (50 samples)
Mixed (50 samples)
None (50 samples)… See the full description on the dataset page: https://huggingface.co/datasets/kenza-ily/icdar_disco.ICDAR2015_OCR
META
https://github.com/open-mmlab/mmocr/blob/main/dataset_zoo/icdar2015/metafile.yml
Name: 'Incidental Scene Text IC15'
Paper:
Title: ICDAR 2015 Competition on Robust Reading
URL: https://rrc.cvc.uab.es/files/short_rrc_2015.pdf
Venue: ICDAR
Year: '2015'
BibTeX: '@inproceedings{karatzas2015icdar,
title={ICDAR 2015 competition on robust reading},
author={Karatzas, Dimosthenis and Gomez-Bigorda, Lluis and Nicolaou, Anguelos and Ghosh, Suman and Bagdanov, Andrew and… See the full description on the dataset page: https://huggingface.co/datasets/MiXaiLL76/ICDAR2015_OCR.ICDAR-2013-Table-Competition-Corrected
ICDAR2013 Table Competition Corrected
This dataset is originally from the ICDAR 2013 Table Competition, but with manual corrections made to the annotations in 2023.
About the original dataset
The original dataset was released as part of the ICDAR 2013 Table Competition.
It can be downloaded here but as of August 2023 accessing the files returns a 403 Forbidden error.
Original license
There is no known license for the original dataset, but the data is commonly… See the full description on the dataset page: https://huggingface.co/datasets/bsmock/ICDAR-2013-Table-Competition-Corrected.
