CoolFace
Modelpublic

xuanthanh318/bge-m3-int4-asym-g32-openvino

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes169downloads
Model Card

BGE-M3 OpenVINO INT4 ASYM g32

Locally compressed derivative of BAAI/bge-m3 for dense embedding inference with OpenVINO. NNCF settings: INT4_ASYM, group size 32, ratio 1.0, scale estimation disabled. NNCF retained protected backup layers as INT8 ASYM.

LoCoMo (English, dense retrieval) macro average: nDCG@10 45.994%, Recall@10 60.878%, MRR@10 44.845%, Hit@10 70.640%. Tested with OpenVINO 2026.3 on Intel Arc 140V GPU. Embedding dimension is 1024 and maximum sequence length is 8192.

Usage

python
import numpy as np
import openvino as ov
from huggingface_hub import snapshot_download
from transformers import AutoTokenizer

repo_id = "YOUR_USERNAME/bge-m3-int4-asym-g32-openvino"
path = snapshot_download(repo_id)
tokenizer = AutoTokenizer.from_pretrained(path)
compiled = ov.Core().compile_model(
    f"{path}/openvino/openvino_model.xml",
    "GPU",  # use "CPU" when needed
)

tokens = tokenizer(
    ["xin chào, đây là phép thử"],
    padding=True,
    truncation=True,
    max_length=8192,
    return_tensors="np",
)
embedding = np.asarray(
    compiled({
        "input_ids": tokens["input_ids"],
        "attention_mask": tokens["attention_mask"],
    })[compiled.output("sentence_embedding")],
    dtype=np.float32,
)
embedding /= np.maximum(np.linalg.norm(embedding, axis=1, keepdims=True), 1e-12)
print(embedding.shape)  # (1, 1024)

This artifact exposes the dense sentence_embedding output. It does not expose BGE-M3 sparse or ColBERT retrieval modes. The original model is MIT licensed; see the BAAI/bge-m3 model card for attribution and citation.