CoolFace
Modelpublic

sinjab/ms-marco-MiniLM-L2-v2-F16-GGUF

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes8downloads
Model Card

ms-marco-MiniLM-L2-v2-F16-GGUF

This model was converted to GGUF format from cross-encoder/ms-marco-MiniLM-L-2-v2 using llama.cpp via the ggml.ai's GGUF-my-repo space.

Refer to the original model card for more details on the model.

Model Information

Quantization Details

This is a F16 quantization of the original model:

  • —F16: Full 16-bit floating point - highest quality, largest size
  • —Q8_0: 8-bit quantization - high quality, good balance
  • —Q4_K_M: 4-bit quantization with medium quality - smaller size, faster inference

Usage

This model can be used with llama.cpp and other GGUF-compatible inference engines.

bash
# Example using llama.cpp
./llama-rerank -m ms-marco-MiniLM-L2-v2-F16.gguf

Model Files

QuantizationUse Case
F16Maximum quality, largest size
Q8_0High quality, good balance of size/performance
Q4KMGood quality, smallest size, fastest inference

Citation

If you use this model, please cite the original model:

bibtex
# See original model card for citation information

License

This model inherits the license from the original model. Please refer to the original model card for license details.

Acknowledgements