Yonatanhaile2026/tigrinya-trocrhandwritten
TrOCR-Handwritten for Tigrinya OCR
  ![Language: Tigrinya]() ![Script: Ge'ez]() 
Tigrinya TrOCR — Handwritten Variant
Adapting TrOCR for Printed Tigrinya Text Recognition: Word-Aware Loss Weighting for Cross-Script Transfer Learning
A fine-tuned TrOCR model for printed Tigrinya line-level text recognition. This is the handwritten pre-training variant, fine-tuned from microsoft/trocr-base-handwritten using vocabulary extension and Word-Aware Loss Weighting to resolve word-boundary failures caused by BPE space-marker conventions.
Model Details
Performance
Evaluated on a held-out test set of 5,000 synthetic Tigrinya text-line images.
Bootstrap 95% Confidence Intervals (1,000 iterations, TrOCR-Printed)
Bootstrap intervals were computed on the TrOCR-Printed variant; see the printed model card for details.
Comparison (same dataset and split)
Training Details
How to Use
from transformers import VisionEncoderDecoderModel, TrOCRProcessor
from PIL import Image
processor = TrOCRProcessor.from_pretrained("Yonatanhaile2026/tigrinya-trocrhandwritten")
model = VisionEncoderDecoderModel.from_pretrained("Yonatanhaile2026/tigrinya-trocrhandwritten")
Load your text-line image
image = Image.open("your_tigrinya_text_line.png").convert("RGB")
pixel_values = processor(images=image, return_tensors="pt").pixel_values
generated_ids = model.generate(pixel_values, num_beams=5, max_length=128)
prediction = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(prediction)
Intended Use
Suitable for:
- Tigrinya OCR research on synthetic or clean text-line images
- Baseline comparison against printed-specific and CTC-based OCR models
- Research on cross-script transfer learning and BPE tokenizer adaptation
Not suitable for:
- Production OCR without validation on real scanned or handwritten documents
- Scenarios where the printed variant would be preferable
- Documents with heavy degradation, low resolution, or non-text noise
Limitations
- Trained and evaluated exclusively on synthetic printed data from a single domain (newspaper text lines)
- Performance on real-world scanned or genuinely handwritten Tigrinya documents is not validated
- Underperforms the printed TrOCR variant on this synthetic printed corpus
- Results reflect a single training run on one hardware configuration
Related Resources
- Printed variant: `Yonatanhaile2026/tigrinya-trocr-printed`
- Code repository: github.com/YoHa2024NKU/Tigrinya_TrOCR_Printed
- Dataset: GLOCR — Harvard Dataverse
Citation
If you use this model, please cite the associated paper and repository:
@misc{medhanie2026adaptingtrocrprintedtigrinya,
title={Adapting TrOCR for Printed Tigrinya Text Recognition: Word-Aware Loss Weighting for Cross-Script Transfer Learning},
author={Yonatan Haile Medhanie and Yuanhua Ni},
year={2026},
eprint={2604.20813},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.20813},
}
License
MIT