sneedium/pixelplanetocr
IterVM: Iterative Vision Modeling Module for Scene Text Recognition
The official code of IterNet.
We propose IterVM, an iterative approach for visual feature extraction which can significantly improve scene text recognition accuracy. IterVM repeatedly uses the high-level visual feature extracted at the previous iteration to enhance the multi-level features extracted at the subsequent iteration.
Runtime Environment
pip install -r requirements.txtNote: fastai==1.0.60 is required.
Datasets
<details> <summary>Training datasets (Click to expand) </summary>
- MJSynth (MJ):
- Use
tools/create_lmdb_dataset.pyto convert images into LMDB dataset - LMDB dataset BaiduNetdisk(passwd:n23k)
- SynthText (ST):
- Use
tools/crop_by_word_bb.pyto crop images from original SynthText dataset, and convert images into LMDB dataset bytools/create_lmdb_dataset.py - LMDB dataset BaiduNetdisk(passwd:n23k)
- WikiText103, which is only used for pre-trainig language models:
- Use
notebooks/prepare_wikitext103.ipynbto convert text into CSV format. - CSV dataset BaiduNetdisk(passwd:dk01) </details>
<details> <summary>Evaluation datasets (Click to expand) </summary>
- Evaluation datasets, LMDB datasets can be downloaded from BaiduNetdisk(passwd:1dbv), GoogleDrive.
- ICDAR 2013 (IC13)
- ICDAR 2015 (IC15)
- IIIT5K Words (IIIT)
- Street View Text (SVT)
- Street View Text-Perspective (SVTP)
- CUTE80 (CUTE) </details>
<details> <summary>The structure of data directory (Click to expand) </summary>
- The structure of
datadirectory is
data
├── charset_36.txt
├── evaluation
│ ├── CUTE80
│ ├── IC13_857
│ ├── IC15_1811
│ ├── IIIT5k_3000
│ ├── SVT
│ └── SVTP
├── training
│ ├── MJ
│ │ ├── MJ_test
│ │ ├── MJ_train
│ │ └── MJ_valid
│ └── ST
├── WikiText-103.csv
└── WikiText-103_eval_d1.csv</details>
Pretrained Models
Get the pretrained models from GoogleDrive. Performances of the pretrained models are summaried as follows:
Training
- Pre-train vision model
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 python main.py --config=configs/pretrain_vm.yaml- Pre-train language model
CUDA_VISIBLE_DEVICES=0,1,2,3 python main.py --config=configs/pretrain_language_model.yaml- Train IterNet
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 python main.py --config=configs/train_iternet.yamlNote:
- You can set the
checkpointpath for vision model (vm) and language model separately for specific pretrained model, or set toNoneto train from scratch
Evaluation
CUDA_VISIBLE_DEVICES=0 python main.py --config=configs/train_iternet.yaml --phase test --image_onlyAdditional flags:
--checkpoint /path/to/checkpointset the path of evaluation model--test_root /path/to/datasetset the path of evaluation dataset--model_eval [alignment|vision]which sub-model to evaluate--image_onlydisable dumping visualization of attention masks
Run Demo
python demo.py --config=configs/train_iternet.yaml --input=figures/demoAdditional flags:
--config /path/to/configset the path of configuration file--input /path/to/image-directoryset the path of image directory or wildcard path, e.g,--input='figs/test/*.png'--checkpoint /path/to/checkpointset the path of trained model--cuda [-1|0|1|2|3...]set the cuda id, by default -1 is set and stands for cpu--model_eval [alignment|vision]which sub-model to use--image_onlydisable dumping visualization of attention masks
Citation
If you find our method useful for your reserach, please cite
@article{chu2022itervm,
title={IterVM: Iterative Vision Modeling Module for Scene Text Recognition},
author={Chu, Xiaojie and Wang, Yongtao},
journal={arXiv preprint arXiv:2204.02630},
year={2022}
}License
The project is only free for academic research purposes, but needs authorization for commerce. For commerce permission, please contact wyt@pku.edu.cn.
Acknowledgements
This project is based on ABINet. Thanks for their great works.
