sinjab/bge-reranker-base-Q4_K_M-GGUF
040
bge-reranker-base-Q4KM-GGUF
This model was converted to GGUF format from BAAI/bge-reranker-base using llama.cpp via the ggml.ai's GGUF-my-repo space.
Refer to the original model card for more details on the model.
Model Information
- Base Model: BAAI/bge-reranker-base
- Quantization: Q4KM
- Format: GGUF (GPT-Generated Unified Format)
- Converted with: llama.cpp
Quantization Details
This is a Q4_K_M quantization of the original model:
- F16: Full 16-bit floating point - highest quality, largest size
- Q8_0: 8-bit quantization - high quality, good balance
- Q4_K_M: 4-bit quantization with medium quality - smaller size, faster inference
Usage
This model can be used with llama.cpp and other GGUF-compatible inference engines.
# Example using llama.cpp
./llama-rerank -m bge-reranker-base-Q4_K_M.ggufModel Files
Citation
If you use this model, please cite the original model:
# See original model card for citation informationLicense
This model inherits the license from the original model. Please refer to the original model card for license details.
Acknowledgements
- Original model by the authors of BAAI/bge-reranker-base
- GGUF conversion via llama.cpp by ggml.ai
- Converted and uploaded by sinjab
