CoolFace
Modelpublic

sowilow/gemma-4-e2b-it-DGX-Spark-GGUF

sourceHugging Facegemmaupdated 6mo agoView on Hugging Face
0likes140downloads
Model Card

πŸš€ v0.1.6: Real-time Metrics & Blackwell-Optimized Docker (Recommended)

This model is fully compatible with the [DGX-Spark-llama.cpp-Bench](https://github.com/sowilow/DGX-Spark-llama.cpp-Bench). Experience the state-of-the-art inference engine optimized for NVIDIA Blackwell (DGX Spark) hardware.

🌟 Key Features (v0.1.6)

  • β€”Real-time Performance Metrics: Now visualizes Input TPS and Output TPS during streaming.
  • β€”Improved Reasoning UI: Seamlessly renders and stabilizes the model's Chain-of-Thought (CoT).
  • β€”Blackwell Optimization: Native support for ARM64/SM121 and CUDA 13.0 FP4.

🐳 Quick Start

bash
# Pull the latest optimized image
docker pull ghcr.io/sowilow/dgx-spark-llama.cpp-bench:v0.1.6

For more details, visit our GitHub Repository.


πŸš€ v0.1.6: μ‹€μ‹œκ°„ μ§€ν‘œ 및 Blackwell μ΅œμ ν™” 도컀 (ꢌμž₯)

이 λͺ¨λΈμ€ [DGX-Spark-llama.cpp-Bench](https://github.com/sowilow/DGX-Spark-llama.cpp-Bench) μ‹œμŠ€ν…œμ— μ΅œμ ν™”λ˜μ–΄ μžˆμŠ΅λ‹ˆλ‹€. NVIDIA Blackwell (DGX Spark) ν•˜λ“œμ›¨μ–΄μ˜ μ„±λŠ₯을 μ΅œλŒ€λ‘œ ν™œμš©ν•˜μ„Έμš”.

🌟 μ£Όμš” νŠΉμ§• (v0.1.6)

  • β€”μ‹€μ‹œκ°„ μ„±λŠ₯ μ§€ν‘œ μ‹œκ°ν™”: 슀트리밍 쀑 Input TPS 및 Output TPSλ₯Ό μ‹€μ‹œκ°„μœΌλ‘œ ν‘œμ‹œν•©λ‹ˆλ‹€.
  • β€”μ§€λŠ₯ν˜• μΆ”λ‘  UI 고도화: λͺ¨λΈμ˜ μƒκ°ν•˜λŠ” κ³Όμ •(CoT)을 더 μ•ˆμ •μ μœΌλ‘œ λ Œλ”λ§ν•©λ‹ˆλ‹€.
  • β€”Blackwell μ΅œμ ν™”: ARM64/SM121 μ•„ν‚€ν…μ²˜ 및 CUDA 13.0 FP4 가속 지원.

🐳 μ‹€ν–‰ 방법

bash
# μ΅œμ‹  μ΅œμ ν™” 이미지 λ‚΄λ €λ°›κΈ°
docker pull ghcr.io/sowilow/dgx-spark-llama.cpp-bench:v0.1.6

μƒμ„Έν•œ μ‚¬μš©λ²•μ€ GitHub 리포지토리λ₯Ό μ°Έμ‘°ν•˜μ„Έμš”.



πŸš€ v0.1.5: Real-time Metrics & Blackwell-Optimized Docker (Recommended)

This model is fully compatible with the [DGX-Spark-llama.cpp-Bench](https://github.com/sowilow/DGX-Spark-llama.cpp-Bench). Experience the state-of-the-art inference engine optimized for NVIDIA Blackwell (DGX Spark) hardware.

🌟 Key Features (v0.1.5)

  • β€”Real-time Performance Metrics: Now visualizes Input TPS and Output TPS during streaming.
  • β€”Improved Reasoning UI: Seamlessly renders and stabilizes the model's Chain-of-Thought (CoT).
  • β€”Blackwell Optimization: Native support for ARM64/SM121 and CUDA 13.0 FP4.

🐳 Quick Start

bash
# Pull the latest optimized image
docker pull ghcr.io/sowilow/dgx-spark-llama.cpp-bench:v0.1.5

For more details, visit our GitHub Repository.


πŸš€ v0.1.5: μ‹€μ‹œκ°„ μ§€ν‘œ 및 Blackwell μ΅œμ ν™” 도컀 (ꢌμž₯)

이 λͺ¨λΈμ€ [DGX-Spark-llama.cpp-Bench](https://github.com/sowilow/DGX-Spark-llama.cpp-Bench) μ‹œμŠ€ν…œμ— μ΅œμ ν™”λ˜μ–΄ μžˆμŠ΅λ‹ˆλ‹€. NVIDIA Blackwell (DGX Spark) ν•˜λ“œμ›¨μ–΄μ˜ μ„±λŠ₯을 μ΅œλŒ€λ‘œ ν™œμš©ν•˜μ„Έμš”.

🌟 μ£Όμš” νŠΉμ§• (v0.1.5)

  • β€”μ‹€μ‹œκ°„ μ„±λŠ₯ μ§€ν‘œ μ‹œκ°ν™”: 슀트리밍 쀑 Input TPS 및 Output TPSλ₯Ό μ‹€μ‹œκ°„μœΌλ‘œ ν‘œμ‹œν•©λ‹ˆλ‹€.
  • β€”μ§€λŠ₯ν˜• μΆ”λ‘  UI 고도화: λͺ¨λΈμ˜ μƒκ°ν•˜λŠ” κ³Όμ •(CoT)을 더 μ•ˆμ •μ μœΌλ‘œ λ Œλ”λ§ν•©λ‹ˆλ‹€.
  • β€”Blackwell μ΅œμ ν™”: ARM64/SM121 μ•„ν‚€ν…μ²˜ 및 CUDA 13.0 FP4 가속 지원.

🐳 μ‹€ν–‰ 방법

bash
# μ΅œμ‹  μ΅œμ ν™” 이미지 λ‚΄λ €λ°›κΈ°
docker pull ghcr.io/sowilow/dgx-spark-llama.cpp-bench:v0.1.5

μƒμ„Έν•œ μ‚¬μš©λ²•μ€ GitHub 리포지토리λ₯Ό μ°Έμ‘°ν•˜μ„Έμš”.



πŸš€ v0.1.4: Quick Start with Blackwell-Optimized Docker (Recommended)

This model is fully compatible with the [DGX-Spark-llama.cpp-Bench](https://github.com/sowilow/DGX-Spark-llama.cpp-Bench). Experience the best performance on NVIDIA Blackwell (DGX Spark) hardware with our optimized inference engine.

🌟 Key Features (v0.1.4)

  • β€”Blackwell Optimized: Native support for ARM64/SM121 and CUDA 13.0 FP4.
  • β€”Intelligent Reasoning UI: Automatic extraction and visualization of reasoning processes (CoT).
  • β€”One-Click Deployment: Standardized environment via GHCR Docker image.

🐳 How to Run

bash
# Pull the latest optimized image
docker pull ghcr.io/sowilow/dgx-spark-llama.cpp-bench:v0.1.4

# Follow the instructions in our repo to serve this model
# GitHub: https://github.com/sowilow/DGX-Spark-llama.cpp-Bench

πŸš€ v0.1.4: Blackwell μ΅œμ ν™” 도컀 ν€΅μŠ€νƒ€νŠΈ (ꢌμž₯)

이 λͺ¨λΈμ€ [DGX-Spark-llama.cpp-Bench](https://github.com/sowilow/DGX-Spark-llama.cpp-Bench) μ‹œμŠ€ν…œμ— μ΅œμ ν™”λ˜μ–΄ μžˆμŠ΅λ‹ˆλ‹€. NVIDIA Blackwell (DGX Spark) ν•˜λ“œμ›¨μ–΄μ˜ μ„±λŠ₯을 μ΅œλŒ€λ‘œ ν™œμš©ν•˜λŠ” μ΅œμ ν™”λœ μΆ”λ‘  엔진을 κ²½ν—˜ν•΄ λ³΄μ„Έμš”.

🌟 μ£Όμš” νŠΉμ§• (v0.1.4)

  • β€”Blackwell μ΅œμ ν™”: ARM64/SM121 μ•„ν‚€ν…μ²˜ 및 CUDA 13.0 FP4 ν•˜λ“œμ›¨μ–΄ 가속 지원.
  • β€”μ§€λŠ₯ν˜• μΆ”λ‘  UI: λͺ¨λΈμ˜ μƒκ°ν•˜λŠ” κ³Όμ •(CoT)을 μžλ™μœΌλ‘œ κ°μ§€ν•˜κ³  μ‹œκ°ν™”ν•©λ‹ˆλ‹€.
  • β€”κ°„νŽΈν•œ 배포: GHCR 도컀 이미지λ₯Ό 톡해 ν™˜κ²½ μ„€μ • 없이 μ¦‰μ‹œ μ‹€ν–‰ κ°€λŠ₯ν•©λ‹ˆλ‹€.

🐳 μ‹€ν–‰ 방법

bash
# μ΅œμ‹  μ΅œμ ν™” 이미지 λ‚΄λ €λ°›κΈ°
docker pull ghcr.io/sowilow/dgx-spark-llama.cpp-bench:v0.1.4

μƒμ„Έν•œ μ‚¬μš©λ²•μ€ GitHub 리포지토리λ₯Ό μ°Έμ‘°ν•˜μ„Έμš”.



πŸš€ Quick Start with Docker (Recommended)

You can easily run this model using the DGX-Spark-llama.cpp-Bench inference engine. It's pre-configured for high-performance inference on NVIDIA hardware (especially Blackwell/DGX Spark).

1. Pull the Docker Image

bash
docker pull ghcr.io/sowilow/dgx-spark-llama.cpp-bench:latest

2. Run the Inference Server

For detailed configuration and usage, visit the GitHub Repository.


gemma-4-e2b-it-GGUF

This repository contains GGUF-quantized weights for Gemma-4-E2B-it, specifically optimized for NVIDIA Blackwell (DGX Spark) hardware.

πŸš€ Key Features

  • β€”Hardware Optimized: Built with CUDA 13.0 and SM121 (Blackwell) native acceleration.
  • β€”Quantization: Q4KM (4-bit unified quantization) for ultra-low latency vision tasks.
  • β€”Base Model Integration: Linked directly to the original google/gemma-4-E2B-it.

βš–οΈ License & Attribution

This model is a quantized version of the original google/gemma-4-E2B-it and is subject to the Gemma License Agreement.

πŸ“‚ Files Included

  • β€”gemma-4-e2b-it-q4_k_m.gguf: Main model weights.
  • β€”gemma-4-e2b-vision-mmproj-f16.gguf: Multimodal vision projector (Dimension-matched: n_embd=1536).

Created using [DGX-Spark-llama.cpp-Bench](https://github.com/sowilow/DGX-Spark-llama.cpp-Bench)