datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Persian_Arabic_TextLine_Image_Ocr_MediumPersian_Arabic_TextLine_Image_Ocr_Small
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/mohajesmaeili/Persian_Arabic_TextLine_Image_Ocr_Small.text-line-imagesdata-towerbooks-textlinesTextline-Detection-Dataset
Textline Detection — Ultimate Dataset
A large-scale, multi-source dataset for textline detection using YOLO-format bounding box annotations.
Combines real scene text, document layout images, and synthetically generated Khmer document images.
Dataset Summary
Split
Images
Train
30,658
Val
2,764
Total
33,422
Classes
ID
Name
Description
0
text_line
A line of Khmer or mixed-script text
1
image
An embedded image/figure region within… See the full description on the dataset page: https://huggingface.co/datasets/Darayut/Textline-Detection-Dataset.MMU-OCR-21-Urdu-TextLineskhmer-textline-dataset
Synthetic Khmer Document Text-Line Detection Dataset
Synthetic dataset for single-class text-line detection on Cambodian
(Khmer) official documents — press releases, ministry letters, formal memos.
Generated with a procedural Pillow-based pipeline featuring:
8 layout templates (standard, letter, announcement, report, sparse,
two-column, memo, plain)
12 page sizes from A5 to A4-landscape
Variable margins, font sizes, line spacing, and indentation
Photometric augmentations… See the full description on the dataset page: https://huggingface.co/datasets/Darayut/khmer-textline-dataset.khmer-textline-dataset_v2
Synthetic Khmer Document Text & Logo Detection Dataset
Synthetic dataset for dual-class text-line and graphical element detection on Cambodian
(Khmer) official documents — press releases, ministry letters, formal memos, and ID cards.
Generated with a procedural Pillow-based pipeline featuring:
8 layout templates (standard, letter, announcement, report, sparse,
two-column, memo, plain)
Procedural Asset Injection (stamps, seals, and logos with perfect bounding boxes)
12 page sizes… See the full description on the dataset page: https://huggingface.co/datasets/Darayut/khmer-textline-dataset_v2.
