CoolFace
Modelpublic

luluw/bge-reranker-v2-m3-eng-nep-16k-trimmed-GGUF

sourceHugging Faceapache-2.0updated 25d agoView on Hugging Face
0likes457downloads
Model Card

Model Card: bge-reranker-v2-m3-eng-nep-16k-trimmed-GGUF

This model card documents the GGUF conversions of the vocabulary-trimmed model `luluw/bge-reranker-v2-m3-eng-nep-16k-trimmed`.

Model Details

  • —Model Name: bge-reranker-v2-m3-eng-nep-16k-trimmed-GGUF
  • —Original (trimmed) Model: `luluw/bge-reranker-v2-m3-eng-nep-16k-trimmed`
  • —Base Model: `BAAI/bge-reranker-v2-m3`
  • —Architecture: XLM-RoBERTa (cross-encoder / sequence classification)
  • —Task: Reranking (outputs a single relevance logit for a (query, passage) pair)
  • —Languages: English + Nepali (vocabulary focused)
  • —Description: Vocabulary-trimmed version of BAAI/bge-reranker-v2-m3. The original multilingual vocabulary (250k+ tokens) was reduced to 16,384 tokens optimized for English and Nepali. No additional fine-tuning was performed — only the input embedding matrix was rebuilt by copying the original embedding rows for the retained tokens. The classification head remains unchanged.

Vocabulary Trimming Summary

MetricOriginalTrimmed
Vocab size (tokenizer)250,00216,384
Model config vocab_size250,00216,384
Model config num_labels11
Max position embeddings8,1948,194 (unchanged)
Embedding parameters256,002,04816,777,216
Total parameters567,755,777328,530,945
Parameters saved—239,224,832 (42.1% smaller)

How the vocabulary was trimmed:

  1. 1.Token frequencies counted on real English + Nepali text (from lbourdois/fineweb-2-trimming).
  2. 2.Special tokens + the first 1,000 original IDs always retained.
  3. 3.Remaining budget filled with highest-frequency English and Nepali tokens (50/50 weighting).
  4. 4.Input embedding matrix rebuilt by copying original rows for kept tokens.
  5. 5.An old_id → new_id mapping (vocab_mapping.json) is provided so the original XLM-R tokenizer can still be used for subword splitting.

Converted Formats & Sizes

FilenameQuant typeFile SizeDescription
bge-reranker-v2-m3-eng-nep-16k-trimmed-fp16.ggufFP16~644 MBFull precision (FP16)
bge-reranker-v2-m3-eng-nep-16k-trimmed-Q8_0.ggufQ8_0~358 MB8-bit quantization
bge-reranker-v2-m3-eng-nep-16k-trimmed-Q6_K.ggufQ6_K~284 MB6-bit K-quantization
bge-reranker-v2-m3-eng-nep-16k-trimmed-Q5_K_M.ggufQ5KM~254 MB5-bit K-quantization (medium)
bge-reranker-v2-m3-eng-nep-16k-trimmed-Q5_K_S.ggufQ5KS~246 MB5-bit K-quantization (small)
bge-reranker-v2-m3-eng-nep-16k-trimmed-Q4_K_M.ggufQ4KM~225 MB4-bit K-quantization (medium)

Usage Notes

  • —These GGUF files are intended for use with llama.cpp (or compatible backends that support XLM-RoBERTa / cross-encoder models).
  • —Because the vocabulary was remapped, you must apply the provided old_to_new mapping when converting token IDs. Do not feed raw original tokenizer IDs directly into the trimmed model.
  • —Recommended pair encoding format remains: <s> query </s></s> passage </s>.
  • —The model outputs a single logit (relevance score). Apply sigmoid if a 0–1 score is desired.
  • —Very rare domain-specific tokens may map to <unk>.

License

The original model is licensed under the MIT License. All converted GGUF files inherit the same license.

Citation

bibtex
@misc{bge_m3,
  title={BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation},
  author={Chen, Jianlv and Xiao, Shitao and Zhang, Peitian and Luo, Kun and Lian, Defu and Liu, Zheng},
  year={2024},
  eprint={2402.03216},
  archivePrefix={arXiv},
  primaryClass={cs.CL}
}