CoolFace
Modelpublic

kernelpool/LongCat-2.0-3bit-UVMAX

sourceHugging Facemitupdated 3mo agoView on Hugging Face
1likes370downloads
Model Card

kernelpool/LongCat-2.0-3bit-UVMAX

Mixed-precision (UVMAX) quantization of meituan-longcat/LongCat-2.0, converted with mlx-lm.

Revision note: originally converted from the FP8 release (meituan-longcat/LongCat-2.0-FP8), the current revision is re-converted from the bf16 master checkpoint.

What is UVMAX?

UVMAX is a mixed-precision scheme: bit widths are assigned per tensor class from measured round-trip quantization error, rather than uniformly. All classes use group size 64.

Tensor classBitsParametersSizeShare
Expert FFNs31.47 T598.5 GiB88.0%
N-gram embedding tables3135 B55.0 GiB8.1%
Attention, dense MLPs631.4 B22.9 GiB3.4%
Embeddings, lm_head62.8 B2.1 GiB0.3%
DSA indexer, MoE routers80.6 B0.6 GiB0.1%
Norms, correction biases (unquantized)0.9 GiB0.1%

Use with mlx

This model requires LongCat-2.0 support from mlx-lm PR #1464, which has not yet been merged. Until it is included in an mlx-lm release, install mlx-lm from the PR branch:

bash
pip install git+https://github.com/ml-explore/mlx-lm.git@refs/pull/1464/head
python
from mlx_lm import load, generate

model, tokenizer = load("kernelpool/LongCat-2.0-3bit-UVMAX")

prompt = "hello"

if tokenizer.chat_template is not None:
    messages = [{"role": "user", "content": prompt}]
    prompt = tokenizer.apply_chat_template(
        messages, add_generation_prompt=True, return_dict=False,
    )

response = generate(model, tokenizer, prompt=prompt, verbose=True)