CoolFace
Modelpublic

anjohn0077/NEXS-qwen3-32b-russian-t-tech-lora

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes4downloads
Model Card

NEXS Qwen3-32B russian-t-tech LoRA (vLLM-ready)

Rank-128 LoRA adapter (bf16) extracted with mergekit from t-tech/T-pro-it-2.0 against the base model Qwen/Qwen3-32B, then sanitized for vLLM serving.

Sanitization applied

The raw mergekit extraction included full-rank modules_to_save tensors (embed_tokens, lm_head, and norm layers) that vLLM's LoRA runtime does not support. This upload contains only the pure low-rank lora_A/lora_B weights (448 pairs: 64 layers x q/k/v/o/gate/up/down projections), with modules_to_save: null in adapter_config.json. No resize_token_embeddings() call is needed to load this adapter.

Serving with vLLM

bash
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen3-32B \
    --enable-lora \
    --lora-modules russian_t_tech=anjohn0077/NEXS-qwen3-32b-russian-t-tech-lora \
    --port 8000 \
    --max-lora-rank 128 \
    --gpu-memory-utilization 0.85

Evaluation (mmmluru)

VariantAccuracy
Base model0.7517
This LoRA on base (via vLLM)0.7542
Original full fine-tune0.7679

Evaluated with lm-evaluation-harness against a local vLLM OpenAI-compatible endpoint:

bash
lm_eval --model local-completions \
    --model_args model=russian_t_tech,base_url=http://localhost:8000/v1/completions,tokenizer=Qwen/Qwen3-32B,num_concurrent=10 \
    --tasks m_mmlu_ru \
    --output_path results/vllm_russian_t_tech