datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai_handwriting_dataset
Thai Handwriting Dataset
This dataset combines two major Thai handwriting datasets:
BEST 2019 Thai Handwriting Recognition dataset (train-0000.parquet)
Thai Handwritten Free Dataset by Wang (train-0001.parquet onwards)
Maintainer
kobkrit@iapp.co.th
Dataset Description
BEST 2019 Dataset
Contains handwritten Thai text images along with their ground truth transcriptions. The images have been processed and standardized for machine learning tasks.… See the full description on the dataset page: https://huggingface.co/datasets/iapp/thai_handwriting_dataset.TreeOil_Painting_ScientificJourney_Thailand_CaseStudy🧪 Tree Oil Painting: A Scientific Journey – Thailand Case Study
This dataset documents a rare and detailed forensic investigation of a mysterious 19th-century oil painting, known as The Tree Oil Painting, using scientific methods and AI-assisted analysis. Compiled in Thailand between 2015 and 2025, this work represents a grassroots effort to validate the painting’s origins through physical evidence, pigment mapping, synchrotron spectroscopy, and historical comparison.
🧩 Overview
Title: Tree… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/TreeOil_Painting_ScientificJourney_Thailand_CaseStudy.ThaiOCRBench
ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
ThaiOCRBench is the first comprehensive benchmark for evaluating vision-language models (VLMs) on Thai text-rich visual understanding tasks.Inspired by OCRBench v2, it contains 2,808 human-annotated samples across 13 diverse tasks, including table parsing, chart understanding, full-page OCR, key information extraction, and visual question answering.
The benchmark enables standardized zero-shot… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ThaiOCRBench.MMMU-Thai
MMMU Thai (MMMU Benchmark Translated to Thai)
MMMU Thai is a dataset for evaluating multimodal models on massive multi-discipline tasks requiring college-level knowledge and deliberate reasoning. This dataset is translated from MMMU (A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI) into Thai.
Dataset Details
MMMU Thai consists of 11,500 meticulously collected multimodal questions from college exams, quizzes, and textbooks… See the full description on the dataset page: https://huggingface.co/datasets/iapp/MMMU-Thai.thai-mscoco-2014-captions
Usage
from datasets import load_dataset
dataset = load_dataset("patomp/thai-mscoco-2014-captions")
dataset
output
DatasetDict({
train: Dataset({
features: ['image', 'filepath', 'sentids', 'filename', 'imgid', 'split', 'sentences_tokens', 'sentences_raw', 'sentences_sentid', 'cocoid', 'th_sentences_raw'],
num_rows: 113287
})
validation: Dataset({
features: ['image', 'filepath', 'sentids', 'filename', 'imgid', 'split', 'sentences_tokens'… See the full description on the dataset page: https://huggingface.co/datasets/patomp/thai-mscoco-2014-captions.thai-ocr-evaluation
Thai OCR Evaluation Dataset
Dataset Description
The Thai OCR Evaluation Dataset is designed for evaluating Optical Character Recognition (OCR) models across various domains. It includes images and textual data derived from various open-source websites.
This dataset aims to provide a comprehensive evaluation resource for researchers and developers working on OCR systems, particularly in Thai language processing.
Data Fields
Each sample in the dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-ocr-evaluation.TreeOil_Painting_ScientificResearch_SortedImageIndex_Thailand2015_2018
Tree Oil Painting – Scientific Research Index (Thailand, 2015–2018)
This dataset documents the scientific and personal journey to investigate the authenticity and artistic origin of The Tree Oil Painting, a mysterious work suspected to be from the late 19th century. The dataset includes structured image records, lab results, motion analysis videos, and historical documents collected between 2015 and 2018 in Thailand.
🌱 Background
In 2015, after visually detecting… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/TreeOil_Painting_ScientificResearch_SortedImageIndex_Thailand2015_2018.thai_handwriting_trio
Thai Handwritten Dataset
This Thai handwritten dataset is curated from three sources:
Ancient scripts [1]
General sentences [2]
Syllables [3]
Dataset Composition
Each canvas contains:
1 ancient script sample
2–5 general sentence samples
4–8 syllable samples
All data were used exactly once, except for the syllable data. Of the 268,056 available syllable samples, only 20,717 were used, balanced across 320 words.
Each canvas has a resolution of 0.4 - 1M px.… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/thai_handwriting_trio.RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts
Dataset Attribution
The original dataset is available on Kaggle.
This dataset has been curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work.
Please cite the original authors if you use this dataset.
Citation
@INPROCEEDINGS{8978005,
author={Chamchong, Rapeeporn and Gao, Wei and McDonnell, Mark D.},
booktitle={2019 International Conference on Document Analysis and Recognition (ICDAR)}… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts.ThaiIDCardSynt
Dataset Details
Dataset Description
Curated by: Matichon Maneegard
Shared by [optional]: Matichon Maneegard
Language(s) (NLP): image-to-text
License: apache-2.0
Dataset Sources [optional]
The dataset was entirely synthetic. It does not contain real information or pertain to any specific person.
Uses
Direct Use
Using for tranning OCR or Multimodal.
Dataset Structure
This dataset contains 98 x 6 = 588 samples, and the… See the full description on the dataset page: https://huggingface.co/datasets/Float16-cloud/ThaiIDCardSynt.thai-handwriting-textdisjoint-v1
Thai Handwriting - Text-Disjoint Split v1
Frozen train/test split of iapp/thai_handwriting_dataset built for the Phase 3
promotion gate of sivakorn-su/typhoon-ocr-7b-thai-handwriting-lora-v1.
Grouped by normalized gold text; every image of a text lands in the same split,
so test texts never appear in train (measures reading, not memorisation).
Test stratified by length bucket with the long bucket over-represented.
CPE-OPH test images are excluded from train by image hash so… See the full description on the dataset page: https://huggingface.co/datasets/sivakorn-su/thai-handwriting-textdisjoint-v1.appen-thai-docdataset_info:
features:
name: image
dtype: image
name: file_name
dtype: string
name: file_type
dtype: string
Original Dataset
The original dataset is available on Kaggle:
https://www.kaggle.com/datasets/appenlimited/ocr-image-data-for-thai-documents
Disclaimer
This dataset is curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work.
Please cite the original author if you use this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/appen-thai-doc.thairocrthai-ocr-text-images
Thai OCR Text Images (Test)
Description
Thai OCR Text Images is a dataset of Thai text images designed for Optical Character Recognition (OCR) and Vision-Language Model (VLM) training.
Each sample consists of:
A cropped image containing Thai text.
The corresponding ground truth text.
The images are generated from PDF documents and are intended to improve Thai text recognition performance in machine learning models.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/pepder/thai-ocr-text-images.thai_handwritten_datasetsThai_Insurance_Docs_OCRThai_e-Government_Procurement_OCRiapp_thai_handwriting_converted_datasetthai_famous_people_images_dataset
Thai Famous People Image Dataset
Dataset Description
The Thai Famous People Image Dataset is a collection of images and descriptions of famous Thai personalities. This dataset is designed to provide a comprehensive resource for researchers, developers, and enthusiasts interested in Thai culture, history, and notable figures. The data was extracted from the Thai Wikipedia dump in September 2024, ensuring up-to-date and relevant information.
Maintainer
Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/iapp/thai_famous_people_images_dataset.thai-road-ocrcoco_caption-thai-ipu24-train-sample10k
Thai Image Captioning Dataset (Samples 10k)
This dataset is a Thai image captioning corpus created to generate natural and human-like image captions in Thai.
It includes a curated subset of 10,000 samples randomly selected from the original training data, providing a smaller yet representative dataset for efficient experimentation.
Dataset Description
The dataset comprises high-quality image captions in Thai, designed for training and evaluating image captioning models.… See the full description on the dataset page: https://huggingface.co/datasets/saksornr/coco_caption-thai-ipu24-train-sample10k.fashion-dataset-thai
fashion-dataset-thai
Thai-localized fashion product dataset: 44,072 product images with metadata fields translated to Thai (gender, category, sub-category, article type, base colour, season, usage). Based on the Fashion Product Images dataset (Kaggle).
Format
Field
Description
id
Product id
year
Year
productDisplayName
Product name
image
Product image
gender_th, masterCategory_th, subCategory_th, articleType_th, baseColour_th, season_th… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/fashion-dataset-thai.product-thai-114k
Product Thai Dataset
Dataset Description
The Product Thai Dataset is a comprehensive collection of Thai product information, including images and ratings, designed for e-commerce and product analysis tasks specific to the Thai market. This dataset contains 113,677 examples, providing a rich source of information for researchers and developers working on product recommendation systems, image-based product search, and e-commerce analytics in the context of Thai products.… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/product-thai-114k.ThaiDataGovernance-datasetthaitrocr-eval-dataset-beta
ThaiTROCR Evaluation Dataset
Dataset Description
The ThaiTROCR Evaluation Dataset is designed for evaluating Optical Character Recognition (OCR) models across various domains. It includes images and textual data derived from various open-source websites.
This dataset aims to provide a comprehensive evaluation resource for researchers and developers working on OCR systems, particularly in Thai language processing.
Data Fields
Each sample in the dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/suchut/thaitrocr-eval-dataset-beta.thai-ocr-test
Thai OCR Evaluation Dataset
Dataset Description
The Thai OCR Evaluation Dataset is designed for evaluating Optical Character Recognition (OCR) models across various domains. It includes images and textual data derived from various open-source websites.
This dataset aims to provide a comprehensive evaluation resource for researchers and developers working on OCR systems, particularly in Thai language processing.
Data Fields
Each sample in the dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Buni111i/thai-ocr-test.Thai_Insurance_Docs_OCR-calib-chandra2Banknote_thaicharts-json-thaiLLAVA_CoT_o1_Instruct_Thai
