onnx-community/minilm-student-L6_uniform_distilled-ONNX
014
minilm-student-L6uniformdistilled (ONNX)
This is an ONNX version of gomyk/minilm-student-L6_uniform_distilled. It was automatically converted and uploaded using this Hugging Face Space.
Usage with Transformers.js
See the pipeline documentation for sentence-similarity: https://huggingface.co/docs/transformers.js/api/pipelines
L6uniformdistilled (Distilled)
Lightweight sentence encoder created from sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 via layer pruning + vocabulary pruning + knowledge distillation.
Model Details
Architecture
==============================================================
TEACHER: MiniLM-L12 → STUDENT: 6L / 38,775 vocab
==============================================================
TEACHER STUDENT
─────────────────────────── ───────────────────────────
┌─────────────────────────┐ ┌─────────────────────────┐
│ Input Tokens │ │ Input Tokens │
└────────────┬────────────┘ └────────────┬────────────┘
│ │
┌────────────┴────────────┐ ┌────────────┴────────────┐
│ Embeddings │ │ Embeddings (pruned) │
│ vocab: 250,002 │ │ vocab: 38,775 │
│ dim: 384 │ │ dim: 384 │
└────────────┬────────────┘ └────────────┬────────────┘
│ │
┌─────────────────────────┐ ┌─────────────────────────┐
│ Layer 0 │ ──► │ Layer 0 ← L0 │
├─────────────────────────┤ ├─────────────────────────┤
│ Layer 1 │ ╳ │ │
├─────────────────────────┤ ├─────────────────────────┤
│ Layer 2 │ ──► │ Layer 1 ← L2 │
├─────────────────────────┤ ├─────────────────────────┤
│ Layer 3 │ ╳ │ │
├─────────────────────────┤ ├─────────────────────────┤
│ Layer 4 │ ──► │ Layer 2 ← L4 │
├─────────────────────────┤ ├─────────────────────────┤
│ Layer 5 │ ╳ │ │
├ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─┤ │ │
│ Layer 6 │ ╳ │ │
├─────────────────────────┤ ├─────────────────────────┤
│ Layer 7 │ ──► │ Layer 3 ← L7 │
├─────────────────────────┤ ├─────────────────────────┤
│ Layer 8 │ ╳ │ │
├─────────────────────────┤ ├─────────────────────────┤
│ Layer 9 │ ──► │ Layer 4 ← L9 │
├─────────────────────────┤ ├─────────────────────────┤
│ Layer 10 │ ╳ │ │
├─────────────────────────┤ ├─────────────────────────┤
│ Layer 11 │ ──► │ Layer 5 ← L11 │
└────────────┬────────────┘ └────────────┬────────────┘
│ │
┌────────────┴────────────┐ ┌────────────┴────────────┐
│ Mean Pooling │ │ Mean Pooling │
│ → 384d embedding │ │ → 384d embedding │
└─────────────────────────┘ └─────────────────────────┘
Size: 448.0MB (FP32) → 98.1MB (FP32)
Params: 117,451,392 → 25,714,176
Reduction: 78.1%
==============================================================Quick Start
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("L6_uniform_distilled", trust_remote_code=True)
sentences = [
"Hello, how are you?",
"안녕하세요",
"Bonjour, comment allez-vous?",
]
embeddings = model.encode(sentences)
print(embeddings.shape) # (3, 384)MTEB Evaluation Results
Overall Average: 56.93%
Classification
Clustering
STS
Distillation Impact
Training
Stage 1: Layer Pruning
- Teacher:
sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2(12 layers, 384d) - Selected layers:
[0, 2, 4, 7, 9, 11](6 layers, evenly spaced (general-purpose)) - Vocabulary pruning applied
Stage 2: Knowledge Distillation
- Method: MSE + Cosine Similarity loss
- Data: MTEB Classification/Clustering/STS task datasets
- Optimizer: AdamW (lr=2e-5, weight_decay=0.01)
- Schedule: Cosine annealing over 3 epochs
Supported Languages (18)
ko, en, ja, zh, es, fr, de, pt, it, ru, ar, hi, th, vi, id, tr, nl, pl
