CoolFace
Modelpublic

samheym/GerColBERT

sourceHugging Faceupdated 1mo agoView on Hugging Face
1likes1kdownloads
Model Card

Model Overview

GerColBERT is a ColBERT-based retrieval model trained on German text. It is designed for efficient late interaction-based retrieval while maintaining high-quality ranking performance. Training Configuration

  • Base Model: deepset/gbert-base
  • Training Dataset: samheym/ger-dpr-collection
  • Dataset: 10% of randomly selected triples from the final dataset
  • Vector Length: 128
  • Maximum Document Length: 256 Tokens
  • Batch Size: 50
  • Training Steps: 80,000
  • Gradient Accumulation: 1 step
  • Learning Rate: 5 × 10⁻⁶
  • Optimizer: AdamW
  • In-Batch Negatives: Included

Usage

Sentence Transformers

This model can be used with Sentence Transformers as a multi-vector (ColBERT-style late interaction) retriever via the MultiVectorEncoder:

bash
pip install "sentence-transformers>=6.0.0"
python
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("samheym/GerColBERT")

query = "Welcher Planet ist als der Rote Planet bekannt?"
documents = [
    "Venus wird wegen ihrer ähnlichen Größe und Nähe oft als Zwilling der Erde bezeichnet.",
    "Mars, bekannt für sein rötliches Aussehen, wird oft als der Rote Planet bezeichnet.",
    "Jupiter, der größte Planet des Sonnensystems, hat einen markanten roten Fleck.",
    "Saturn, berühmt für seine Ringe, wird manchmal mit dem Roten Planeten verwechselt.",
]

query_embeddings = model.encode_query(query)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings[0].shape)
# (32, 128) (18, 128)

# MaxSim late-interaction scoring (higher is more relevant)
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[13.8994, 27.9157, 21.6549, 22.7871]])

PyLate

First install the PyLate library:

bash
pip install -U pylate

Retrieval

PyLate provides a streamlined interface to index and retrieve documents using ColBERT models. The index leverages the Voyager HNSW index to efficiently handle document embeddings and enable fast retrieval.

python
from pylate import indexes, models, retrieve

# Step 1: Load the ColBERT model
model = models.ColBERT(
    model_name_or_path=samheym/GerColBERT,
)

<!--

Citation

BibTeX

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->