CoolFace
Modelpublic

URajinda/Qwen2.5-MM-1.5B-Base

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
0likes30downloads
Model Card

Qwen2.5-MM-1.5B-Base

မြန်မာစာအတွက် ပြုပြင်ပြီး Qwen2.5-1.5B Base Model

Model Description

This is a Myanmar language adapted version of Qwen2.5-1.5B. The model has been extended with Myanmar vocabulary for better tokenization efficiency and language understanding. Qwen2.5-MM-1.5B-Base is a Myanmar-centric base language model built on top of the Qwen 2.5 1.5B architecture. This model is a milestone in the "ShweYon" project, focusing on improving the efficiency of Myanmar script processing through a custom tokenizer.

Key Features

  • —Extended vocabulary with 500+ Myanmar tokens
  • —75-85% reduction in token count for Myanmar text
  • —4x training speed for Myanmar Language!!!
  • —Compatible with original Qwen2.5 architecture
  • —Ready for Myanmar language fine-tuning
  • —Still Need pre-training for Myanmar Language

Technical Details

  • —Base Model: Qwen2.5-1.5B
  • —Extended Vocabulary Size: 152,165 tokens
  • —Original Vocabulary Size: 151,665 tokens
  • —Myanmar Tokens Added: 500 tokens
  • —Model Parameters: 1.54B

Optimization Tip

Don't just fine-tune. Redefine your tokenizer for language efficiency. A small, dense vocabulary (500 tokens) can be 4x faster and more stable than generic ones.

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

# ၁။ Device ကို အရင်သတ်မှတ်မယ် (GPU ရှိရင် cuda၊ မရှိရင် cpu)
device = "cuda" if torch.cuda.is_available() else "cpu"

model_name = "URajinda/Qwen2.5-MM-1.5B-Base"

# ၂။ Model ကို load လုပ်ပြီး GPU ဆီ ပို့မယ် (.to(device) ထည့်လိုက်တာပါ)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16).to(device)
tokenizer = AutoTokenizer.from_pretrained(model_name)

# ၃။ စာရိုက်မယ်
prompt = "မြန်မာနိုင်ငံတွင်"

# ၄။ Input တွေကိုလည်း GPU ဆီ ပို့မယ် (.to(device) ထည့်လိုက်တာပါ)
inputs = tokenizer(prompt, return_tensors="pt").to(device)

# ၅။ အဖြေထုတ်မယ်
outputs = model.generate(
    **inputs,
    max_new_tokens=100,
    temperature=0.7,
    do_sample=True
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))