xuanthanh318/bge-m3-int4-asym-g64-openvino
0172
BGE-M3 OpenVINO INT4 ASYM g64
Locally compressed derivative of BAAI/bge-m3 for dense embedding inference with OpenVINO. NNCF settings: INT4_ASYM, group size 64, ratio 1.0, scale estimation disabled. NNCF retained protected backup layers as INT8 ASYM.
LoCoMo (English, dense retrieval) macro average: nDCG@10 44.644%, Recall@10 59.425%, MRR@10 43.496%, Hit@10 68.601%. Tested with OpenVINO 2026.3 on Intel Arc 140V GPU. Embedding dimension is 1024 and maximum sequence length is 8192.
Usage
import numpy as np
import openvino as ov
from huggingface_hub import snapshot_download
from transformers import AutoTokenizer
repo_id = "YOUR_USERNAME/bge-m3-int4-asym-g64-openvino"
path = snapshot_download(repo_id)
tokenizer = AutoTokenizer.from_pretrained(path)
compiled = ov.Core().compile_model(
f"{path}/openvino/openvino_model.xml",
"GPU", # use "CPU" when needed
)
tokens = tokenizer(
["xin chào, đây là phép thử"],
padding=True,
truncation=True,
max_length=8192,
return_tensors="np",
)
embedding = np.asarray(
compiled({
"input_ids": tokens["input_ids"],
"attention_mask": tokens["attention_mask"],
})[compiled.output("sentence_embedding")],
dtype=np.float32,
)
embedding /= np.maximum(np.linalg.norm(embedding, axis=1, keepdims=True), 1e-12)
print(embedding.shape) # (1, 1024)This artifact exposes the dense sentence_embedding output. It does not expose BGE-M3 sparse or ColBERT retrieval modes. The original model is MIT licensed; see the BAAI/bge-m3 model card for attribution and citation.
