CoolFace
Modelpublic

Roflmax/bge-m3-legal-ru-cocktail-40-60

sourceHugging Facemitupdated 11mo agoView on Hugging Face
0likes623downloads
Model Card

BGE-M3 Legal RU Cocktail (40/60)

🏆 Top-performing Russian legal document embedding model with 91.79% Recall@5

This model is created using LM-Cocktail weight interpolation technique, combining two state-of-the-art Russian legal embedding models:

Synergistic Effect: The cocktail model outperforms both parent models through optimal weight combination!

Model Details

Model Description

  • —Model Type: Sentence Transformer (LM-Cocktail)
  • —Base Models: BGE-M3 fine-tuned on Russian legal documents
  • —Maximum Sequence Length: 512 tokens (optimized for speed)
  • —Output Dimensionality: 1024 dimensions
  • —Similarity Function: Cosine Similarity
  • —Language: Russian
  • —Domain: Legal documents (court decisions, federal laws, regional legislation)
  • —License: MIT

Key Features

  • —✅ Best Recall@5 among all tested models (91.79%)
  • —✅ No prefix required (inherited from bge-m3-legal-ru-updata)
  • —✅ Balanced performance across all legal document types
  • —✅ Production-ready for Russian legal semantic search

Performance

Benchmark Results (dataset2: 7,187 test examples)

MetricScoreRank
Recall@176.66%#1
Recall@591.79%🥇 #1
Recall@1094.85%#1

Performance by Dataset Type

DatasetRecall@1Recall@5Recall@10Description
court_law66.01%86.47%91.08%Court decisions and rulings
other_law90.44%95.80%96.97%Federal laws and codes
reg_law75.52%93.09%96.49%Regional legislation
Average76.66%91.79%94.85%Across all domains

Comparison with Parent Models

ModelRecall@5Improvement
Cocktail 40/60 (this model)91.79%Baseline
bge-m3-russian-legal91.43%+0.36%
bge-m3-legal-ru-updata91.28%+0.51%

The cocktail demonstrates synergistic effect - it outperforms both parent models!

Usage

Installation

bash
pip install -U sentence-transformers

Basic Usage

python
from sentence_transformers import SentenceTransformer

# Load the model
model = SentenceTransformer("Roflmax/bge-m3-legal-ru-cocktail-40-60")
model.max_seq_length = 512  # Optimized for speed

# Example: Semantic search in legal documents
query = "Какое наказание предусмотрено за управление транспортным средством в состоянии опьянения?"

documents = [
    "Статья 264.1 УК РФ. Нарушение правил дорожного движения лицом, подвергнутым административному наказанию...",
    "КоАП РФ Статья 12.8. Управление транспортным средством водителем, находящимся в состоянии опьянения...",
    "Статья 228 УК РФ. Незаконные приобретение, хранение, перевозка, изготовление..."
]

# Encode
query_embedding = model.encode(query, normalize_embeddings=True)
doc_embeddings = model.encode(documents, normalize_embeddings=True)

# Calculate similarity
from sklearn.metrics.pairwise import cosine_similarity
similarities = cosine_similarity([query_embedding], doc_embeddings)[0]

# Get top results
top_indices = similarities.argsort()[::-1]
for idx in top_indices:
    print(f"Score: {similarities[idx]:.4f} | {documents[idx][:100]}...")

Batch Processing

python
from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("Roflmax/bge-m3-legal-ru-cocktail-40-60")
model.max_seq_length = 512

# Batch encode documents
documents = [
    "Первый документ...",
    "Второй документ...",
    # ... more documents
]

# Process in batches for efficiency
embeddings = model.encode(
    documents,
    batch_size=32,
    normalize_embeddings=True,
    show_progress_bar=True
)

print(f"Generated {len(embeddings)} embeddings of dimension {embeddings.shape[1]}")

Semantic Search Pipeline

python
from sentence_transformers import SentenceTransformer, util

model = SentenceTransformer("Roflmax/bge-m3-legal-ru-cocktail-40-60")

# Your corpus
corpus = [
    "Документ 1: содержание...",
    "Документ 2: содержание...",
    # ... more documents
]

# Encode corpus once
corpus_embeddings = model.encode(corpus, convert_to_tensor=True, normalize_embeddings=True)

# Query
query = "Ваш поисковый запрос"
query_embedding = model.encode(query, convert_to_tensor=True, normalize_embeddings=True)

# Search
hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=5)[0]

# Display results
for hit in hits:
    print(f"Score: {hit['score']:.4f} | {corpus[hit['corpus_id']][:100]}...")

Important Notes

No Prefix Required

Unlike some BGE models, this model does NOT require query/passage prefixes. Simply encode your text directly:

python
# ✅ Correct - no prefix needed
embedding = model.encode("Ваш текст")

# ❌ Not needed
embedding = model.encode("Represent this sentence for searching relevant passages: Ваш текст")

Sequence Length

The model is optimized for 512 tokens:

  • —Fast inference speed
  • —Minimal quality loss (< 1% documents truncated)
  • —Ideal for most legal document fragments

For longer documents, consider chunking:

python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("Roflmax/bge-m3-legal-ru-cocktail-40-60")
model.max_seq_length = 512

# Split long document into chunks
def chunk_text(text, max_length=2000):
    # Simple character-based chunking
    return [text[i:i+max_length] for i in range(0, len(text), max_length)]

long_document = "Очень длинный документ..."
chunks = chunk_text(long_document)
chunk_embeddings = model.encode(chunks, normalize_embeddings=True)

# Use average embedding for the whole document
import numpy as np
document_embedding = np.mean(chunk_embeddings, axis=0)

Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'XLMRobertaModel'})
  (1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True})
  (2): Normalize()
)

Training Details

LM-Cocktail Mixing

This model was created using LM-Cocktail technique (paper), which performs linear weight interpolation between two fine-tuned models:

W_cocktail = 0.4 × W_russian-legal + 0.6 × W_updata

Parent Models

  1. 1.bge-m3-russian-legal (40% weight)
  2. 2.26K training examples (70% deduplicated)
  3. 3.3 epochs, BS=64
  4. 4.Specialized for Russian legal domain
  5. 5.Requires query prefix
  1. 1.bge-m3-legal-ru-updata (60% weight)
  2. 2.54K training examples
  3. 3.5 epochs, BS=128
  4. 4.No prefix required
  5. 5.Strong on regional legislation

Why 40/60 Ratio?

The 40/60 weight ratio was selected based on comprehensive benchmarking:

  • —Tested 5 ratios: 30/70, 40/60, 50/50, 60/40, 70/30
  • —40/60 and 50/50 achieved identical top Recall@5 (91.79%)
  • —40/60 selected for slightly better Recall@1 (76.66% vs 76.47%)

Use Cases

1. Legal Document Search

Find relevant laws, court decisions, and regulations based on natural language queries.

2. Document Classification

Classify legal documents by type, jurisdiction, or topic using embedding similarity.

3. Duplicate Detection

Identify similar or duplicate legal documents in large corpora.

4. Question Answering

Build semantic search systems for legal Q&A applications.

5. Document Clustering

Group related legal documents for analysis and organization.

Limitations

  • —Language: Optimized for Russian only
  • —Domain: Legal documents (may underperform on other domains)
  • —Sequence Length: 512 tokens (longer documents require chunking)
  • —Recency: Training data cutoff unknown (inherited from parent models)

Citation

If you use this model, please cite:

bibtex
@misc{bge-m3-legal-ru-cocktail,
  title={BGE-M3 Legal RU Cocktail: LM-Cocktail Weight Interpolation for Russian Legal Embeddings},
  author={Roflmax},
  year={2025},
  publisher={HuggingFace},
  howpublished={\url{https://huggingface.co/Roflmax/bge-m3-legal-ru-cocktail-40-60}}
}

Also cite the parent models:

And the LM-Cocktail paper:

bibtex
@article{jiang2023lmcocktail,
  title={LM-Cocktail: Resilient Tuning of Language Models via Model Merging},
  author={Jiang, Shitao and others},
  journal={arXiv preprint arXiv:2311.13534},
  year={2023}
}

Framework Versions

  • —Python: 3.12.3
  • —Sentence Transformers: 5.1.2
  • —Transformers: 4.57.1
  • —PyTorch: 2.8.0+cu128
  • —LM-Cocktail: 0.0.4

License

MIT License - free for commercial and non-commercial use.

Contact

  • —Author: Roflmax
  • —HuggingFace: Roflmax
  • —Issues: Report issues on the model page

Note: This model achieves state-of-the-art performance on Russian legal document retrieval tasks. For best results, use with normalize_embeddings=True and cosine similarity.