Roflmax/bge-m3-legal-ru-cocktail-40-60
BGE-M3 Legal RU Cocktail (40/60)
🏆 Top-performing Russian legal document embedding model with 91.79% Recall@5
This model is created using LM-Cocktail weight interpolation technique, combining two state-of-the-art Russian legal embedding models:
- 40% Roflmax/bge-m3-russian-legal (91.43% Recall@5)
- 60% Roflmax/bge-m3-legal-ru-updata (91.28% Recall@5)
Synergistic Effect: The cocktail model outperforms both parent models through optimal weight combination!
Model Details
Model Description
- Model Type: Sentence Transformer (LM-Cocktail)
- Base Models: BGE-M3 fine-tuned on Russian legal documents
- Maximum Sequence Length: 512 tokens (optimized for speed)
- Output Dimensionality: 1024 dimensions
- Similarity Function: Cosine Similarity
- Language: Russian
- Domain: Legal documents (court decisions, federal laws, regional legislation)
- License: MIT
Key Features
- ✅ Best Recall@5 among all tested models (91.79%)
- ✅ No prefix required (inherited from bge-m3-legal-ru-updata)
- ✅ Balanced performance across all legal document types
- ✅ Production-ready for Russian legal semantic search
Performance
Benchmark Results (dataset2: 7,187 test examples)
Performance by Dataset Type
Comparison with Parent Models
The cocktail demonstrates synergistic effect - it outperforms both parent models!
Usage
Installation
pip install -U sentence-transformersBasic Usage
from sentence_transformers import SentenceTransformer
# Load the model
model = SentenceTransformer("Roflmax/bge-m3-legal-ru-cocktail-40-60")
model.max_seq_length = 512 # Optimized for speed
# Example: Semantic search in legal documents
query = "Какое наказание предусмотрено за управление транспортным средством в состоянии опьянения?"
documents = [
"Статья 264.1 УК РФ. Нарушение правил дорожного движения лицом, подвергнутым административному наказанию...",
"КоАП РФ Статья 12.8. Управление транспортным средством водителем, находящимся в состоянии опьянения...",
"Статья 228 УК РФ. Незаконные приобретение, хранение, перевозка, изготовление..."
]
# Encode
query_embedding = model.encode(query, normalize_embeddings=True)
doc_embeddings = model.encode(documents, normalize_embeddings=True)
# Calculate similarity
from sklearn.metrics.pairwise import cosine_similarity
similarities = cosine_similarity([query_embedding], doc_embeddings)[0]
# Get top results
top_indices = similarities.argsort()[::-1]
for idx in top_indices:
print(f"Score: {similarities[idx]:.4f} | {documents[idx][:100]}...")Batch Processing
from sentence_transformers import SentenceTransformer
import numpy as np
model = SentenceTransformer("Roflmax/bge-m3-legal-ru-cocktail-40-60")
model.max_seq_length = 512
# Batch encode documents
documents = [
"Первый документ...",
"Второй документ...",
# ... more documents
]
# Process in batches for efficiency
embeddings = model.encode(
documents,
batch_size=32,
normalize_embeddings=True,
show_progress_bar=True
)
print(f"Generated {len(embeddings)} embeddings of dimension {embeddings.shape[1]}")Semantic Search Pipeline
from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer("Roflmax/bge-m3-legal-ru-cocktail-40-60")
# Your corpus
corpus = [
"Документ 1: содержание...",
"Документ 2: содержание...",
# ... more documents
]
# Encode corpus once
corpus_embeddings = model.encode(corpus, convert_to_tensor=True, normalize_embeddings=True)
# Query
query = "Ваш поисковый запрос"
query_embedding = model.encode(query, convert_to_tensor=True, normalize_embeddings=True)
# Search
hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=5)[0]
# Display results
for hit in hits:
print(f"Score: {hit['score']:.4f} | {corpus[hit['corpus_id']][:100]}...")Important Notes
No Prefix Required
Unlike some BGE models, this model does NOT require query/passage prefixes. Simply encode your text directly:
# ✅ Correct - no prefix needed
embedding = model.encode("Ваш текст")
# ❌ Not needed
embedding = model.encode("Represent this sentence for searching relevant passages: Ваш текст")Sequence Length
The model is optimized for 512 tokens:
- Fast inference speed
- Minimal quality loss (< 1% documents truncated)
- Ideal for most legal document fragments
For longer documents, consider chunking:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("Roflmax/bge-m3-legal-ru-cocktail-40-60")
model.max_seq_length = 512
# Split long document into chunks
def chunk_text(text, max_length=2000):
# Simple character-based chunking
return [text[i:i+max_length] for i in range(0, len(text), max_length)]
long_document = "Очень длинный документ..."
chunks = chunk_text(long_document)
chunk_embeddings = model.encode(chunks, normalize_embeddings=True)
# Use average embedding for the whole document
import numpy as np
document_embedding = np.mean(chunk_embeddings, axis=0)Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'XLMRobertaModel'})
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True})
(2): Normalize()
)Training Details
LM-Cocktail Mixing
This model was created using LM-Cocktail technique (paper), which performs linear weight interpolation between two fine-tuned models:
W_cocktail = 0.4 × W_russian-legal + 0.6 × W_updataParent Models
- bge-m3-russian-legal (40% weight)
- 26K training examples (70% deduplicated)
- 3 epochs, BS=64
- Specialized for Russian legal domain
- Requires query prefix
- bge-m3-legal-ru-updata (60% weight)
- 54K training examples
- 5 epochs, BS=128
- No prefix required
- Strong on regional legislation
Why 40/60 Ratio?
The 40/60 weight ratio was selected based on comprehensive benchmarking:
- Tested 5 ratios: 30/70, 40/60, 50/50, 60/40, 70/30
- 40/60 and 50/50 achieved identical top Recall@5 (91.79%)
- 40/60 selected for slightly better Recall@1 (76.66% vs 76.47%)
Use Cases
1. Legal Document Search
Find relevant laws, court decisions, and regulations based on natural language queries.
2. Document Classification
Classify legal documents by type, jurisdiction, or topic using embedding similarity.
3. Duplicate Detection
Identify similar or duplicate legal documents in large corpora.
4. Question Answering
Build semantic search systems for legal Q&A applications.
5. Document Clustering
Group related legal documents for analysis and organization.
Limitations
- Language: Optimized for Russian only
- Domain: Legal documents (may underperform on other domains)
- Sequence Length: 512 tokens (longer documents require chunking)
- Recency: Training data cutoff unknown (inherited from parent models)
Citation
If you use this model, please cite:
@misc{bge-m3-legal-ru-cocktail,
title={BGE-M3 Legal RU Cocktail: LM-Cocktail Weight Interpolation for Russian Legal Embeddings},
author={Roflmax},
year={2025},
publisher={HuggingFace},
howpublished={\url{https://huggingface.co/Roflmax/bge-m3-legal-ru-cocktail-40-60}}
}Also cite the parent models:
And the LM-Cocktail paper:
@article{jiang2023lmcocktail,
title={LM-Cocktail: Resilient Tuning of Language Models via Model Merging},
author={Jiang, Shitao and others},
journal={arXiv preprint arXiv:2311.13534},
year={2023}
}Framework Versions
- Python: 3.12.3
- Sentence Transformers: 5.1.2
- Transformers: 4.57.1
- PyTorch: 2.8.0+cu128
- LM-Cocktail: 0.0.4
License
MIT License - free for commercial and non-commercial use.
Contact
- Author: Roflmax
- HuggingFace: Roflmax
- Issues: Report issues on the model page
Note: This model achieves state-of-the-art performance on Russian legal document retrieval tasks. For best results, use with normalize_embeddings=True and cosine similarity.
