CoolFace
Modelpublic

amjd-ai/Qwen2.5-3B-KV-Compressed-50

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes107downloads
Model Card

Qwen2.5-3B KV-Compressed (50% Memory Reduction)

An optimized, fully-fused version of Qwen2.5-3B featuring real-time 50% Key-Value (KV) Cache Compression using Bipartite Cosine Similarity token merging inside the self-attention mechanism.

Benchmark & Key Achievements

  • —50.0% Direct KV Memory Reduction: Halves KV cache footprint during long-context generation.
  • —Bilingual & Coding Competence: Preserves full conversational reasoning in Arabic, English, and Python.
  • —Zero Generation Latency Overhead: Highly efficient runtime execution with full gradient alignment.

Model Details

  • —Base Architecture: Qwen2.5-3B
  • —Compression Mechanism: Dynamic KV Chunk Merging ($K$ & $V$ projection alignment)
  • —Parameters: 3.09B