NeuronUz/NeuronAI-2B
NeuronAI-2B
NeuronAI-2B is an Uzbek-first, bilingual assistant model built from Qwen3.5-2B-Base. It combines an Uzbek tokenizer retrofit, continued pretraining, annealing, and assistant-only supervised fine-tuning. The published weights are fully merged—no LoRA adapter is needed.
License: Apache License 2.0. Commercial and non-commercial use are permitted under the license terms. This differs from the NeuronAI-4B release, which is licensed for non-commercial use.
Quick start
Install a recent Transformers build with Qwen3.5 support:
pip install -U "transformers>=5.1" accelerate torchimport torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "NeuronUz/NeuronAI-2B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
dtype=torch.bfloat16,
device_map={"": 0},
).eval()
messages = [
{"role": "system", "content": "Siz foydali va aniq AI yordamchisiz."},
{"role": "user", "content": "Alisher Navoiy haqida qisqacha aytib bering."},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
enable_thinking=False,
return_tensors="pt",
return_dict=True,
).to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=1024,
do_sample=True,
temperature=0.7,
top_p=0.8,
top_k=20,
min_p=0.0,
repetition_penalty=1.0,
use_cache=True,
)
reply = tokenizer.decode(
output[0, inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
).strip()
print(reply)This is the recommended quality-oriented preset for general assistant use: non-thinking mode with Qwen3.5's instruct sampling settings. Greedy decoding can cause repetition and lower response quality; reserve do_sample=False for deterministic evaluation or classification. The generation metadata already registers <|im_end|> and <|endoftext|> as end-of-sequence tokens. Keep the combined prompt and output within the validated 4,096-token serving limit.
Serve with vLLM
pip install -U vllm
vllm serve NeuronUz/NeuronAI-2B \
--dtype bfloat16 \
--max-model-len 4096 \
--tensor-parallel-size 1 \
--generation-config vllm \
--default-chat-template-kwargs '{"enable_thinking":false}' \
--language-model-only \
--enable-prefix-caching \
--mamba-block-size 16 \
--mamba-cache-mode aligncurl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "NeuronUz/NeuronAI-2B",
"messages": [
{"role": "user", "content": "O‘zbekiston haqida uchta fakt ayting."}
],
"max_tokens": 1024,
"temperature": 0.7,
"top_p": 0.8,
"top_k": 20,
"min_p": 0.0,
"presence_penalty": 1.5,
"repetition_penalty": 1.0,
"chat_template_kwargs": {"enable_thinking": false}
}'Benchmarks
All five model result sets below cover the same full eight-task suite. Classification and multiple-choice tasks use accuracy; FLORES+ translation uses COMET. The weighted score is normalized by the 0.95 sum of the published task weights. All eight NeuronAI-2B tasks completed and passed the invalid-output gate.
Alloma runs used the APST apostrophe preprocessing required by their model cards; NeuronAI and stock Qwen did not. The alloma-8B column combines its full model-card-protocol evaluation with separately archived full UzLiB, TUMLU-Uzbek, and MMLU-Uzbek runs. Exact source files, scores, and run IDs are included in `benchmark_results.json`.
Run the benchmarks on your computer
The repository includes a portable Alloma-style benchmark runner. It covers FLORES+ (both directions), Uzbek sentiment, Uzbek news, MMLU English, MMLU Uzbek, and TUMLU-Uzbek.
pip install -r https://huggingface.co/NeuronUz/NeuronAI-2B/resolve/main/benchmark-requirements.txt
wget https://huggingface.co/NeuronUz/NeuronAI-2B/resolve/main/benchmark.py
python benchmark.py --limit 200 --output quick-results.jsonThe quick command uses the same seed on 200 examples per dataset. Run all public examples and add COMET with:
pip install unbabel-comet
python benchmark.py --limit 0 --comet --output full-results.jsonRun one task when you only need a short check:
python benchmark.py --tasks mmlu-uz --limit 200 --output mmlu-uz.json
python benchmark.py --tasks flores --limit 200 --output flores.json--limit 0 means the full dataset. Only full runs are comparable with the table above; 200-example quick runs are sanity checks. COMET downloads the Unbabel/wmt22-comet-da evaluator and needs additional disk/RAM.
Uzbek tokenizer efficiency
The tokenizer is an in-place, primarily Latin-script Uzbek retrofit rather than a vocabulary extension. The initial 20,000-document figure was measured on training-source uz-crawl, so we replaced it with a larger corpus-stratified test: 118,832 held-out-source documents plus a separate 100,000-document training-source control. Documents were selected with deterministic SHA-256 bottom-k sampling (seed 20260825), exact duplicates were excluded from the selected sample, tiny texts were filtered, and raw source text was tokenized without apostrophe normalization.
Across the two held-out sources combined, the tokenizer uses 35.19% fewer tokens overall and 40.90% fewer tokens on Latin-dominant text, matching its intended Latin-Uzbek focus.
The paired intervals use 5,000 bootstrap replicates over 1,000 deterministic document buckets. OSCAR may still have incidental overlap with other public web corpora and was previously checked in a post-hoc weak-token coverage analysis, but it contributed no tokenizer-training rows. The legal corpus does not appear in the tokenizer or training source manifests and is the cleanest source-and-domain holdout in this test. Full results and script/length breakdowns: `fertility_large_20260825.json` and `fertility_large_20260825.md`.
Fertility measures tokenization efficiency—not model quality or measured decoding speed. The evaluated 2B and 4B custom tokenizer files are byte-identical, as are their evaluated stock-base tokenizer files; SHA-256 fingerprints are recorded in the JSON result.
Training
The mixture is Uzbek-first and includes general assistant conversations, translation, Uzbek language and literature, spelling, classification, math, and English-retention examples. Training data is not distributed with this model repository.
Intended use
Good fits include Uzbek research, education, commercial and non-commercial prototyping, translation experiments, writing assistance, retrieval-augmented generation, and local/offline applications. Users remain responsible for validating the model for their application and complying with the Apache 2.0 license and applicable law.
Limitations
- This is a public-suite-selected checkpoint. The benchmark results are useful for reproducibility and relative comparison, but they are not a locked, independent estimate of real-world generalization.
- LoRA rank, learning rate, batch size, and dropout were not exhaustively swept; the table reports the released run, not globally optimal hyperparameters.
- TUMLU-Uzbek is the weakest reported Uzbek task and should not be treated as solved at 32.57% accuracy.
- The model can hallucinate, repeat biases in its data, or produce unsafe or outdated content. It has not been comprehensively safety-evaluated.
- Do not rely on it without expert review for medical, legal, financial, public safety, or other high-stakes decisions.
- SFT used sequences up to 2,048 tokens; serving at longer inherited context lengths has not been validated here. The published inference examples use 4,096 tokens.
License
NeuronAI-2B is released under the Apache License 2.0. Commercial and non-commercial use, modification, and distribution are permitted subject to its terms. This summary does not replace the license text; see `LICENSE`.
