CoolFace
Modelpublic

stefan-jo/mxbai-edge-colbert-v0-32m-scifact-kmeans-pf1-6

sourceHugging Faceotherupdated 6mo agoView on Hugging Face
0likes16downloads
Model Card

mxbai-edge-colbert-v0-32m SciFact KMeans PF1-6

This is a research checkpoint based on `mixedbread-ai/mxbai-edge-colbert-v0-32m`, fine-tuned on SciFact training data with pooling-aware distillation. During training, document embeddings were pooled with k-means and the pool factor was sampled uniformly from 1 to 6 on each batch.

This released checkpoint corresponds to the selected final SciFact run used for the public artifact release. It is intended primarily for research, reproduction, and further experimentation with pooling-aware late interaction retrieval.

Model Details

  • —Base model: mixedbread-ai/mxbai-edge-colbert-v0-32m
  • —Architecture: ColBERT / late interaction retrieval
  • —Query length: 48
  • —Document length: 300
  • —Embedding dimension: 64
  • —Training setup: k-means pooling, multi-factor PF1-6
  • —Training dataset: stefan-jo/scifact-train-mined-reranker-scores

Intended Use

This model is intended for:

  • —reproducing the paper's SciFact pooling-aware fine-tuning results
  • —experimenting with pooling-aware late interaction retrieval
  • —studying how multi-factor training affects retrieval under document compression

It is not presented as a general-purpose production retriever.

Usage

python
from pylate import models

model = models.ColBERT(
    model_name_or_path="stefan-jo/mxbai-edge-colbert-v0-32m-scifact-kmeans-pf1-6"
)

queries_embeddings = model.encode(
    ["What evidence supports the claim?"],
    is_query=True,
)

documents_embeddings = model.encode(
    ["The abstract provides evidence relevant to the claim."],
    is_query=False,
    pool_factor=4,
    pool_method="kmeans",
    use_sklearn=True,
)

Evaluation Snapshot

The table below is adapted from the paper's NanoBEIR cross-dataset effects table. All runs use k-means pooling at inference and report NDCG@10.

DatasetPFBaselineFT SciFact KMeans PF1-6FT FiQA KMeans PF1-6
SciFact10.8080.8170.802
SciFact20.7650.8130.774
SciFact40.6490.8100.808
SciFact60.6090.7950.758
FiQA10.5260.5230.528
FiQA20.4880.5190.505
FiQA40.4700.4900.513
FiQA60.4310.4590.467
NFCorpus10.3750.3720.369
NFCorpus20.3700.3700.378
NFCorpus40.3420.3630.381
NFCorpus60.3070.3630.369
SCIDOCS10.3960.3930.382
SCIDOCS20.3710.3850.375
SCIDOCS40.3470.3740.387
SCIDOCS60.3320.3800.371
Touché202010.5960.5950.592
Touché202020.5970.6020.601
Touché202040.5650.5940.572
Touché202060.5450.5730.571

For full experiments and additional tables, see the accompanying paper and repository:

Training Data Provenance

The training dataset was built from SciFact train data using a mining-and-reranking pipeline:

  • —hard negative mining with BAAI/bge-small-en-v1.5
  • —teacher scores from BAAI/bge-reranker-v2-gemma
  • —distillation training on mined candidate sets with reranker scores

The released training dataset is available separately as stefan-jo/scifact-train-mined-reranker-scores.

License and Provenance

This model is released under a mixed-source custom notice for caution. The underlying SciFact-based training data combines multiple upstream sources, including SciFact annotations and S2ORC-derived corpus content.

Relevant upstream components:

  • —base model: mixedbread-ai/mxbai-edge-colbert-v0-32m (Apache-2.0)
  • —mining model: BAAI/bge-small-en-v1.5 (MIT)
  • —reranker: BAAI/bge-reranker-v2-gemma (Apache-2.0)
  • —training data source: SciFact / BEIR-style preprocessing
  • —SciFact annotations: CC BY 4.0
  • —SciFact corpus source: S2ORC / ODC-By 1.0

Users should review and comply with the upstream attribution and source terms in addition to this repository notice.