CoolFace
Datasetpublic

openthaigpt/thai-ocr-evaluation

Thai OCR Evaluation Dataset Dataset Description The Thai OCR Evaluation Dataset is designed for evaluating Optical Character Recognition (OCR) models across various domains. It includes images and textual data derived from various open-source websites. This dataset aims to provide a comprehensive evaluation resource for researchers and developers working on OCR systems, particularly in Thai language processing. Data Fields Each sample in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-ocr-evaluation.

sourceHugging Facecc-by-sa-4.0updated 2y agoView on Hugging Face
9likes202downloads
Dataset Card

Thai OCR Evaluation Dataset

Dataset Description

The Thai OCR Evaluation Dataset is designed for evaluating Optical Character Recognition (OCR) models across various domains. It includes images and textual data derived from various open-source websites. This dataset aims to provide a comprehensive evaluation resource for researchers and developers working on OCR systems, particularly in Thai language processing.

Data Fields

Each sample in the dataset contains the following fields:

  • image: Path to the image file.
  • text: Ground truth text extracted from the image.
  • category: The domain/category of the image (e.g., "handwritten", "document", "scene_text").

Usage

To load the dataset, you can use the following code:

python
from datasets import load_dataset

dataset = load_dataset("openthaigpt/thai-ocr-evaluation")

Sponsors

<img src="https://cdn-uploads.huggingface.co/production/uploads/66f6b837fbc158f2846a9108/Cvr5AWLNque5BljUX-C-n.png" alt="Sponsors" width="50%">

Authors

  • Suchut Sapsathien (suchut@outlook.com)
  • Jillaphat Jaroenkantasima (autsadang41@gmail.com)