CoolFace
Modelpublic

sinjab/bge-reranker-large-F16-GGUF

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes33downloads
Model Card

bge-reranker-large-F16-GGUF

This model was converted to GGUF format from BAAI/bge-reranker-large using llama.cpp via the ggml.ai's GGUF-my-repo space.

Refer to the original model card for more details on the model.

Model Information

  • —Base Model: BAAI/bge-reranker-large
  • —Quantization: F16
  • —Format: GGUF (GPT-Generated Unified Format)
  • —Converted with: llama.cpp

Quantization Details

This is a F16 quantization of the original model:

  • —F16: Full 16-bit floating point - highest quality, largest size
  • —Q8_0: 8-bit quantization - high quality, good balance
  • —Q4_K_M: 4-bit quantization with medium quality - smaller size, faster inference

Usage

This model can be used with llama.cpp and other GGUF-compatible inference engines.

bash
# Example using llama.cpp
./llama-rerank -m bge-reranker-large-F16.gguf

Model Files

QuantizationUse Case
F16Maximum quality, largest size
Q8_0High quality, good balance of size/performance
Q4KMGood quality, smallest size, fastest inference

Citation

If you use this model, please cite the original model:

bibtex
# See original model card for citation information

License

This model inherits the license from the original model. Please refer to the original model card for license details.

Acknowledgements