CoolFace
Modelpublic

sahilchachra/hy-mt2-7b-4bit-mlx

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes50downloads
Model Card

hy-mt2-7b-4bit-mlx

Quantized version of tencent/Hy-MT2-7B for Apple Silicon using MLX.

Hy-MT2-7B is Tencent's multilingual translation model covering 40+ languages.

Quantization: Affine integer quantization Precision: 4-bit (~4.5 bits/weight avg) Group size: 64 Disk size: 4042 MB Quantized by: sahilchachra

About this variant

Standard affine (integer) quantization at 4-bit with group size 64. Largest compression ratio — recommended when memory is tight or you want the fastest decode throughput.

Benchmark results

Evaluated on Apple M5 Pro with MLX. Model loaded once; performance and quality measured in a single pass.

Performance

This modelFP16 baseline
Prefill (tok/s)486.87307.33
Decode (tok/s)65.5419.55
Peak memory (GB)4.53215.171
Disk size (MB)404215331

Translation quality (FLORES-200 devtest)

Reported as chrF++ (higher is better). Sample-size noted per pair.

DirectionThis modelFP16 baselinen
engLatn→fraLatn68.3568.7420
engLatn→deuLatn63.8763.2520
engLatn→zhoHans30.3829.420
engLatn→jpnJpan40.9542.2820
engLatn→spaLatn56.9656.720
fraLatn→engLatn68.2867.9920
zhoHans→engLatn57.7857.4320
jpnJpan→engLatn59.0159.220

Avg chrF++: 60.35 vs FP16 60.24 Avg BLEU: 35.86 vs FP16 35.35

Context scaling (decode tok/s)

Context lengthDecode tok/s
~128 tokens64.1
~256 tokens63.8
~512 tokens63.8
~1024 tokens63.2

Usage

Install

bash
pip install mlx-lm

Translate

python
from mlx_lm import load, generate

model, tokenizer = load("sahilchachra/hy-mt2-7b-4bit-mlx")

prompt = (
    "Translate the following text from English to French.\n"
    "English: The early bird catches the worm.\n"
    "French:"
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=128, verbose=True))

Stream

python
from mlx_lm import load, stream_generate

model, tokenizer = load("sahilchachra/hy-mt2-7b-4bit-mlx")
for chunk in stream_generate(model, tokenizer, prompt="Translate \"Hello world\" to Japanese:", max_tokens=64):
    print(chunk.text, end="", flush=True)

All variants in this collection

ModelMethod
sahilchachra/hy-mt2-7b-4bit-mlxAffine int4 (group 64) ← this model
sahilchachra/hy-mt2-7b-8bit-mlxAffine int8 (group 64)

Notes

  • —Requires Apple Silicon (M1 or later) with MLX
  • —Benchmarks run on Apple M5 Pro, 24 GB unified memory
  • —FLORES-200 sample sizes are small — treat chrF/BLEU figures as indicative, not definitive
  • —License: see tencent/Hy-MT2-7B for the original model's license terms

Original model

See tencent/Hy-MT2-7B for full model details, supported languages, and intended use.