datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-ocr-bench
Persian OCR Evaluation Dataset
This benchmark contains paired Persian document images and UTF-8 transcription
targets for OCR evaluation. Each JSONL row references one image under
bench_data/images/ and contains its transcription in the text field.
Schema
image: image path relative to the repository
id: stable image identifier
page: page number, currently 1
type: evaluation item type, currently transcription
text: reference transcription
language: fa
checked:… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian-ocr-bench.persian-ocr-benchmark
Persian OCR Evaluation Dataset
This benchmark contains paired Persian document images and UTF-8 transcription
targets for OCR evaluation. Each JSONL row references one image under
bench_data/images/ and contains its transcription in the text field.
Schema
image: image path relative to the repository
id: stable image identifier
page: page number, currently 1
type: evaluation item type, currently transcription
text: reference transcription
language: fa
checked:… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-ocr-benchmark.OCRBenchV2-DocParsing-UpdatedGTOCRBenchV2-DocParsing-UpdatedGT is an improved ground truth for the document parsing subset of OCRBench V2, created and verified by Tensorlake.
This version is used in Tensorlake’s OCR and document understanding benchmark comparisons.
The dataset is intended for evaluation and research purposes only.For the original benchmark and other subsets, please refer to OCRBench V2
.
