CoolFace
Modelpublic

kernelpool/LongCat-2.0-3bit

sourceHugging Facemitupdated 3mo agoView on Hugging Face
3likes401downloads
Model Card

kernelpool/LongCat-2.0-3bit

3-bit quantization of meituan-longcat/LongCat-2.0, converted with mlx-lm.

Revision note: originally converted from the FP8 release (meituan-longcat/LongCat-2.0-FP8), the current revision is re-converted from the bf16 master checkpoint.

Use with mlx

This model requires LongCat-2.0 support from mlx-lm PR #1464, which has not yet been merged. Until it is included in an mlx-lm release, install mlx-lm from the PR branch:

bash
pip install git+https://github.com/ml-explore/mlx-lm.git@refs/pull/1464/head
python
from mlx_lm import load, generate

model, tokenizer = load("kernelpool/LongCat-2.0-3bit")

prompt = "hello"

if tokenizer.chat_template is not None:
    messages = [{"role": "user", "content": prompt}]
    prompt = tokenizer.apply_chat_template(
        messages, add_generation_prompt=True, return_dict=False,
    )

response = generate(model, tokenizer, prompt=prompt, verbose=True)