CoolFace
Modelpublic

OttomanNLP/Azra-1-Mini-0.8b

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
4likes664downloads
Model Card

Azra 1-Mini (0.8B) — Ottoman Turkish OCR Model

<p align="center"> <img src="banner.png" alt="Azra 1-Mini Banner" width="45%"> </p>

Azra 1-Mini 0.8b is a lightweight, high-performance Vision-Language OCR model specialized in Ottoman Turkish text transcription across both Nesih (printed/calligraphic) and Rika/Riqa (handwritten) scripts.

Despite having only 0.8 billion parameters, Azra 1-Mini achieves state-of-the-art accuracy on Ottoman Turkish OCR tasks, outperforming significantly larger proprietary models.

⚠️ Important Note on Input Resolution & Segmentation: This model has been fine-tuned and optimized specifically for line-level text images (satır bazlı görüntüler). It may not achieve optimal accuracy directly on full-page images without prior text line cropping/segmentation.

📊 Benchmark Results & Performance Comparison

The model was evaluated against leading proprietary Vision-Language models on standard Ottoman Turkish test sets using character accuracy (100% - CER).

1. Nesih Script Test Set (Printed / Calligraphic)

ModelSuccess Rate (%)Rank
Gemini 3.1 Pro82.82%👑 1st
Azra 1-Mini 0.8b80.23%🥈 2nd
Qwen 3.8 Max74.26%🥉 3rd

2. Rika Script Test Set (Handwritten)

ModelSuccess Rate (%)Rank
Azra 1-Mini 0.8b67.08%👑 1st (Winner)
Gemini 3.1 Pro58.46%🥈 2nd
Qwen 3.8 Max54.99%🥉 3rd
🌟 Key Highlight: Azra 1-Mini 0.8b achieves 1st place on the handwritten Rika dataset (67.08%), significantly outperforming both Gemini 3.1 Pro and Qwen 3.8 Max while running efficiently at sub-billion parameter scale.

📷 Qualitative Results & Sample Transcriptions

Below are top qualitative predictions generated by Azra 1-Mini 0.8b from the evaluation test sets:

1. Nesih Script Samples (Printed / Calligraphic)

ImageGround Truth (GT)Model Prediction (Azra 1-Mini)CER
<img src="assets/samples/nesih3263759633eScline_f5695a2d.png" height="35">امّا اری وابدار ونازک اولور هر اعجک زمان غرسیامّا اری وابدار ونازک اولور هر اعجک زمان غرسی0.00%
<img src="assets/samples/nesih3263759704eScline_211ddfad.png" height="35">هلاک ایدر ازایسه علاج ایله خلاص اولورهلاک ایدر ازایسه علاج ایله خلاص اولور0.00%
<img src="assets/samples/nesih3263759800eScline_83cb8596.png" height="35">یافوجی ایچنه دوشرلر اوّل التنه وافرد وکلمش خردلیافوجی ایچنه دوشرلر اوّل التنه وافرد وکلمش خردل0.00%
<img src="assets/samples/nesih3263759772eScline_21130720.png" height="35">اغزی محکم باغلنوب اول بوداق اکلوب یره کوملسه وقتاغزی محکمه باغلنوب اول بوداق اکلوب یره کوملسه وقت2.08%
<img src="assets/samples/nesih3263759629eScline_e9d59dc1.png" height="35">دکمک زماندر دیمش یعنی آیک نقصانی زمانی که اوّلدکک زماندر دیمش یعنی آیک نقصانی زمانی که اوّل2.17%

2. Rika Script Samples (Handwritten)

ImageGround Truth (GT)Model Prediction (Azra 1-Mini)CER
<img src="assets/samples/riqa_dfc88259-fa14-4ae1-b772-f5ea65db8b6c-001.png" height="35">دیمک طلب و دیانت بزدن، دین، شریعت، هدایت اللهدندر و بو هدایت ایکیدیک طلب و دیانت بزدن، دین، شریعت، هدایت اللهدندر۔ و بوهدایت ایکی4.62%
<img src="assets/samples/riqa_dfc88259-fa14-4ae1-b772-f5ea65db8b6c-015.png" height="35">ایتمک دون بنی تنویر ایدن کونشک یارین تنویر ایدهمیهجکنی ادعا ایتمک کبی قانون استقرایی انکاردر۔ایتمک دوند بنی تنویر ایدن کونشک یارین تنویر ایدرمهجیکنی ادعا ایتمک کبی قانون استقرالی انکاردر۔5.38%
<img src="assets/samples/riqa_c25bfa03-dbf2-41e2-8164-d639b87fbede-015.png" height="35">ایمانده نه قدر بیوک بر سعادت و نعمت؛ و نه قدر بیوک بر لذت و راحت بولوندیغنی اڭلامقایمانده نه قدر یوک بر سعادت ونعمت و نه قدر یوک بر لذت و راحت بولوندیغی اشلامم8.54%
<img src="assets/samples/riqa_e2da2c9f-5dc9-4ffa-8241-d13cffc1caef-020.png" height="35">”الله تعالی ابراهیم علیه السلامه وحی ایدوب دیدی که: اسماعیل حقندهکی دعاکی قبول ایتدم و اونی"الله تعالی ابراهیم علمه السلام دحی ایدوب دیدی کی: اسماعیل حقندهکی دعاک قبول ایتدم واولی8.79%
<img src="assets/samples/riqa_81f5dd07-3ee3-4b26-a02a-a589e289f11c-003.png" height="35">بوراده مطلوب اولمامق لازم کلیر، فی الواقع "الصراط المستقیم" نظم جلیلی بزه علی الاطلاقبوراده مطلوب اولاسون لازم کلیر۔ فی الواقع "الصراط المستقیم" نظام جلیلی بزه علی الاطام9.41%

🚀 Usage Guide (transformers)

Below is the standard, native PyTorch & Hugging Face transformers implementation using AutoProcessor and Qwen3_5ForConditionalGeneration:

python
import os
import torch
from PIL import Image
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
from qwen_vl_utils import process_vision_info

# Device & dtype settings
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if device == "cuda" else torch.float32

model_id = "OttomanNLP/Azra-1-Mini-0.8b"

print("[INFO] Loading model and processor...")
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    model_id,
    torch_dtype=dtype,
    device_map="auto" if device == "cuda" else None,
    trust_remote_code=True
)
model.eval()
print("[INFO] Model loaded successfully!")

def extract_text(image_path: str, prompt: str = "Görseldeki Osmanlıca metni transkribe et:") -> str:
    """Extract Ottoman text from a line image"""
    if not os.path.exists(image_path):
        return f"File not found: {image_path}"
        
    image = Image.open(image_path).convert("RGB")
    
    # Adjust dimensions to multiples of 64
    w, h = image.size
    new_w = ((w + 63) // 64) * 64
    new_h = ((h + 63) // 64) * 64
    if (new_w, new_h) != (w, h):
        image = image.resize((new_w, new_h), Image.Resampling.LANCZOS)
    
    messages = [{
        "role": "user",
        "content": [
            {"type": "image", "image": image},
            {"type": "text", "text": prompt}
        ]
    }]
    
    text_input = processor.apply_chat_template(
        messages, tokenize=False, add_generation_prompt=True
    )
    image_inputs, _ = process_vision_info(messages)
    
    inputs = processor(
        text=[text_input],
        images=image_inputs,
        padding=True,
        return_tensors="pt"
    ).to(device)
    
    with torch.inference_mode():
        generated_ids = model.generate(
            **inputs,
            max_new_tokens=512,
            do_sample=False,
            repetition_penalty=1.2,
            no_repeat_ngram_size=3,
            pad_token_id=processor.tokenizer.pad_token_id,
            eos_token_id=processor.tokenizer.eos_token_id,
        )
    
    input_len = inputs.input_ids.shape[1]
    output_text = processor.batch_decode(
        generated_ids[:, input_len:],
        skip_special_tokens=True,
        clean_up_tokenization_spaces=False
    )[0]
    
    return output_text.strip()

if __name__ == "__main__":
    image_path = "sample_line.png"  # Path to line-level image
    text = extract_text(image_path)
    print("📝 Transcribed Text:\n", text)

🏷️ Model Details

  • —Developed by: OttomanNLP
  • —Authors: Gökhan Usta, Oğuz Alpoğlu, Fatih Günaydın
  • —Model Type: Vision-Language Model (VLM) for OCR
  • —Language(s): Ottoman Turkish (Osmanlıca)
  • —Base Architecture: Qwen3.5-Vision
  • —Parameters: ~0.8B
  • —License: Apache-2.0

📚 Citation

If you use this model or dataset in your research, please cite our paper:

bibtex
@article{usta2026cross,
  title={Cross-Lingual Transfer Learning and Autonomous Data Bootstrapping for VLM-Based Ottoman Turkish Handwritten Text Recognition},
  author={Usta, G{\"o}khan and Alpo{\u{g}}lu, O{\u{g}}uz and G{\"u}nayd{\i}n, Fatih},
  journal={Research Square (Preprint)},
  year={2026},
  doi={10.21203/rs.3.rs-10418926/v1},
  note={Under Review at International Journal on Document Analysis and Recognition (IJDAR)}
}

APA:

Usta, G., Alpoğlu, O., & Günaydın, F. (2026). Cross-Lingual Transfer Learning and Autonomous Data Bootstrapping for VLM-Based Ottoman Turkish Handwritten Text Recognition. Research Square Preprint. DOI: 10.21203/rs.3.rs-10418926/v1

📄 License & Attribution

This model is released under the Apache 2.0 License. Free for commercial and research use.