CoolFace
Datasetpublic

sivakorn-su/thai-handwriting-textdisjoint-v1

Thai Handwriting - Text-Disjoint Split v1 Frozen train/test split of iapp/thai_handwriting_dataset built for the Phase 3 promotion gate of sivakorn-su/typhoon-ocr-7b-thai-handwriting-lora-v1. Grouped by normalized gold text; every image of a text lands in the same split, so test texts never appear in train (measures reading, not memorisation). Test stratified by length bucket with the long bucket over-represented. CPE-OPH test images are excluded from train by image hash so… See the full description on the dataset page: https://huggingface.co/datasets/sivakorn-su/thai-handwriting-textdisjoint-v1.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes114downloads
Dataset Card

Thai Handwriting - Text-Disjoint Split v1

Frozen train/test split of iapp/thai_handwriting_dataset built for the Phase 3 promotion gate of sivakorn-su/typhoon-ocr-7b-thai-handwriting-lora-v1.

  • —Grouped by normalized gold text; every image of a text lands in the same split, so test texts never appear in train (measures reading, not memorisation).
  • —Test stratified by length bucket with the long bucket over-represented.
  • —CPE-OPH test images are excluded from train by image hash so legacy comparisons on that set stay valid.
  • —Construction params + leak checks: split_manifest.json.

Apache-2.0, inherited from the source dataset.