datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
neurips_glotocr
GlotOCR-bench
GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages.
Quick links:
🏆 Leaderboard
📝 License
neurips_glotocr
GlotOCR-bench
GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages.
Quick links:
🏆 Leaderboard
📝 License
GlotOCR-bench
GlotOCR-bench
GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages.
Quick links:
📃 Paper
🛠️ Code
📈 Results
🏆 Leaderboard
License
This dataset is released under the GlotOCR Open Evaluation License v1.0 (see LICENSE file for full terms).
The GlotOCR-bench metadata is licensed under CC0-1.0.
The texts used to… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotOCR-bench.GlotOCR-bench-v1.0-results
GlotOCR Bench Results
This dataset contains OCR results from images in cis-lmu/GlotOCR-bench using multiple multilingual OCR models.
Quick links:
📃 Paper
🛠️ Code
📈 Benchmark
Download
git clone https://huggingface.co/datasets/cis-lmu/GlotOCR-bench-v1.0-results
If you want to clone without large files - just their pointers:
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/cis-lmu/GlotOCR-bench-v1.0-results
Processing Details
Source… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotOCR-bench-v1.0-results.
