oddadmix/Emhotob-5M-v2
051
Emhotob-5M-v2
Emhotob is a family of small Arabic language models pretrained from scratch on Arabic web text. This is the 5M rung of the ladder (~5.08M parameters), part of a scaling series ranging from 500K to 25M parameters that all share the same tokenizer, context length, and training recipe. This is the v2 run: trained from scratch on ~5B tokens per model.
⚠️ These are tiny, research-scale models trained on a limited token budget. They are intended for scaling-law experiments, education, and Arabic NLP research — not for production use.
Model details
Training
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "oddadmix/Emhotob-5M-v2"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16)
prompt = "الذكاء الاصطناعي هو"
inputs = tok(prompt, return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=50, do_sample=True, top_p=0.9, temperature=0.8)
print(tok.decode(out[0], skip_special_tokens=True))Limitations
Given its size and limited pretraining budget, Emhotob-5M-v2 has a narrow capability range and will produce factually unreliable and sometimes incoherent text. It has not been instruction-tuned or aligned, and no safety filtering has been applied. Use accordingly.
© SupraLabs 2026 — PROJECT EMHOTOB.
