luluw/bge-reranker-v2-m3-eng-nep-16k-trimmed-GGUF
0457
Model Card: bge-reranker-v2-m3-eng-nep-16k-trimmed-GGUF
This model card documents the GGUF conversions of the vocabulary-trimmed model `luluw/bge-reranker-v2-m3-eng-nep-16k-trimmed`.
Model Details
- Model Name:
bge-reranker-v2-m3-eng-nep-16k-trimmed-GGUF - Original (trimmed) Model: `luluw/bge-reranker-v2-m3-eng-nep-16k-trimmed`
- Base Model: `BAAI/bge-reranker-v2-m3`
- Architecture: XLM-RoBERTa (cross-encoder / sequence classification)
- Task: Reranking (outputs a single relevance logit for a
(query, passage)pair) - Languages: English + Nepali (vocabulary focused)
- Description: Vocabulary-trimmed version of BAAI/bge-reranker-v2-m3. The original multilingual vocabulary (250k+ tokens) was reduced to 16,384 tokens optimized for English and Nepali. No additional fine-tuning was performed — only the input embedding matrix was rebuilt by copying the original embedding rows for the retained tokens. The classification head remains unchanged.
Vocabulary Trimming Summary
How the vocabulary was trimmed:
- Token frequencies counted on real English + Nepali text (from
lbourdois/fineweb-2-trimming). - Special tokens + the first 1,000 original IDs always retained.
- Remaining budget filled with highest-frequency English and Nepali tokens (50/50 weighting).
- Input embedding matrix rebuilt by copying original rows for kept tokens.
- An
old_id → new_idmapping (vocab_mapping.json) is provided so the original XLM-R tokenizer can still be used for subword splitting.
Converted Formats & Sizes
Usage Notes
- These GGUF files are intended for use with
llama.cpp(or compatible backends that support XLM-RoBERTa / cross-encoder models). - Because the vocabulary was remapped, you must apply the provided
old_to_newmapping when converting token IDs. Do not feed raw original tokenizer IDs directly into the trimmed model. - Recommended pair encoding format remains:
<s> query </s></s> passage </s>. - The model outputs a single logit (relevance score). Apply sigmoid if a 0–1 score is desired.
- Very rare domain-specific tokens may map to
<unk>.
License
The original model is licensed under the MIT License. All converted GGUF files inherit the same license.
Citation
@misc{bge_m3,
title={BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation},
author={Chen, Jianlv and Xiao, Shitao and Zhang, Peitian and Luo, Kun and Lian, Defu and Liu, Zheng},
year={2024},
eprint={2402.03216},
archivePrefix={arXiv},
primaryClass={cs.CL}
}