CoolFace
Modelpublic

dolfsai/Qwen3-Embedding-0.6B-vllm-W8A8

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
1likes426downloads
Model Card

prudant/Qwen3-Embedding-0.6B-W8A8

This is a compressed version of Qwen/Qwen3-Embedding-0.6B using llm-compressor with the following scheme: W8A8

Important: You MUST read the following guide for correct usage of this model here Guide

Model Details

  • —Original Model: Qwen/Qwen3-Embedding-0.6B
  • —Quantization Method: GPTQ
  • —Compression Libraries: llm-compressor
  • —Calibration Dataset: ultrachat_200k (1024 samples)
  • —Optimized For: Inference with vLLM
  • —License: same as original model