anjohn0077/NEXS-qwen3-32b-russian-refalmachine-lora
06
NEXS Qwen3-32B russian-refalmachine LoRA (vLLM-ready)
Rank-128 LoRA adapter (bf16) extracted with mergekit from RefalMachine/RuadaptQwen3-32B-Instruct against the base model Qwen/Qwen3-32B, then sanitized for vLLM serving.
Sanitization applied
The raw mergekit extraction included full-rank modules_to_save tensors (embed_tokens, lm_head, and norm layers) that vLLM's LoRA runtime does not support. This upload contains only the pure low-rank lora_A/lora_B weights (448 pairs: 64 layers x q/k/v/o/gate/up/down projections), with modules_to_save: null in adapter_config.json. No resize_token_embeddings() call is needed to load this adapter.
Serving with vLLM
python -m vllm.entrypoints.openai.api_server \
--model Qwen/Qwen3-32B \
--enable-lora \
--lora-modules russian_refalmachine=anjohn0077/NEXS-qwen3-32b-russian-refalmachine-lora \
--port 8000 \
--max-lora-rank 128 \
--gpu-memory-utilization 0.85Evaluation (mmmluru)
Evaluated with lm-evaluation-harness against a local vLLM OpenAI-compatible endpoint:
lm_eval --model local-completions \
--model_args model=russian_refalmachine,base_url=http://localhost:8000/v1/completions,tokenizer=Qwen/Qwen3-32B,num_concurrent=10 \
--tasks m_mmlu_ru \
--output_path results/vllm_russian_refalmachine