sahilchachra/hy-mt2-7b-4bit-mlx
hy-mt2-7b-4bit-mlx
Quantized version of tencent/Hy-MT2-7B for Apple Silicon using MLX.
Hy-MT2-7B is Tencent's multilingual translation model covering 40+ languages.
Quantization: Affine integer quantization Precision: 4-bit (~4.5 bits/weight avg) Group size: 64 Disk size: 4042 MB Quantized by: sahilchachra
About this variant
Standard affine (integer) quantization at 4-bit with group size 64. Largest compression ratio — recommended when memory is tight or you want the fastest decode throughput.
Benchmark results
Evaluated on Apple M5 Pro with MLX. Model loaded once; performance and quality measured in a single pass.
Performance
Translation quality (FLORES-200 devtest)
Reported as chrF++ (higher is better). Sample-size noted per pair.
Avg chrF++: 60.35 vs FP16 60.24 Avg BLEU: 35.86 vs FP16 35.35
Context scaling (decode tok/s)
Usage
Install
pip install mlx-lmTranslate
from mlx_lm import load, generate
model, tokenizer = load("sahilchachra/hy-mt2-7b-4bit-mlx")
prompt = (
"Translate the following text from English to French.\n"
"English: The early bird catches the worm.\n"
"French:"
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=128, verbose=True))Stream
from mlx_lm import load, stream_generate
model, tokenizer = load("sahilchachra/hy-mt2-7b-4bit-mlx")
for chunk in stream_generate(model, tokenizer, prompt="Translate \"Hello world\" to Japanese:", max_tokens=64):
print(chunk.text, end="", flush=True)All variants in this collection
Notes
- Requires Apple Silicon (M1 or later) with MLX
- Benchmarks run on Apple M5 Pro, 24 GB unified memory
- FLORES-200 sample sizes are small — treat chrF/BLEU figures as indicative, not definitive
- License: see tencent/Hy-MT2-7B for the original model's license terms
Original model
See tencent/Hy-MT2-7B for full model details, supported languages, and intended use.
