CoolFace
Modelpublic

NANI-Nithin/Ling-3.0-tiny-GGUF

sourceHugging Faceupdated 18d agoView on Hugging Face
8likes5.6kdownloads
Model Card

Ling-3.0-tiny GGUF (llama.cpp)

![Model Type](https://github.com/ggerganov/llama.cpp) ![Backend](https://github.com/ggerganov/llama.cpp) ![Size](https://huggingface.co/NANI-Nithin/Ling-3.0-tiny-GGUF)

Quantized GGUF files for Ling-3.0-tiny, IBM's lightweight hybrid reasoning MoE model optimized for deployment with llama.cpp. This model delivers strong reasoning and agentic capabilities at low inference cost through an efficient hybrid architecture combining KDA and MLA attention with a sparse MoE FFN.

Model Overview

Ling-3.0-tiny is a 7.9B parameter model with only 1.3B activated parameters per token, designed for efficient local and edge deployment. It features:

  • Efficient Hybrid-Linear Architecture: 3:1 alternating stacking of KDA and MLA (3 KDA layers followed by 1 MLA layer per 4-layer block) with a sparse MoE FFN comprising 128 routed experts
  • Native Hybrid Reasoning and Agentic Capabilities: Supports both fast responses and multi-step reasoning through configurable thinking mode
  • Local and Edge Deployment: Validated on NVIDIA DGX Spark, Apple Silicon MacBook, and Mac mini for capable reasoning without datacenter-class GPUs

Key Capabilities

  • Parameter-Efficient MoE: Only 1.3B of 7.9B parameters activated per token for balanced performance and efficiency
  • Hybrid Attention: Combines KDA (Kimi Delta Attention) and MLA (Multi-Head Latent Attention) for efficient long-context processing
  • Fast Inference: Reaches ~100-105 tokens/s on DGX Spark and 86-90 tokens/s on M4 Pro MacBook with FP8
  • Memory Efficient: ~8.34 GB peak memory usage at 8K context length
  • Thinking Mode: Native chain-of-thought reasoning with per-request configurability

Available Quantizations

FileSizeQualityRecommended Use
Ling-3.0-tiny-BF16.gguf14.72 GBFull precision source. Every quant below is cut from this file.Original model, maximum quality
Ling-3.0-tiny-F16.gguf14.72 GBFull precision source.Alternative full precision
Ling-3.0-tiny-Q8_0.gguf7.83 GBEffectively lossless. Use when disk and RAM are not the constraint.Highest quality, less compression
Ling-3.0-tiny-Q6_K.gguf6.05 GBNear-lossless; the last stop before quality becomes measurable.Balanced quality/size
Ling-3.0-tiny-Q5_K_M.gguf5.25 GBVery good quality, noticeably smaller than Q6_K.Good trade-off
Ling-3.0-tiny-Q5_K_S.gguf5.11 GBSlightly smaller than Q5KM for a slight quality cost.Smaller footprint
Ling-3.0-tiny-Q5_1.gguf5.55 GBLegacy. Prefer Q5KM.Historical compatibility
Ling-3.0-tiny-Q5_0.gguf5.11 GBLegacy. Prefer Q5KM.Historical compatibility
Ling-3.0-tiny-Q4_K_M.gguf4.49 GBThe usual default. Best quality-per-byte for most people.Default choice
Ling-3.0-tiny-Q4_K_S.gguf4.24 GBA little smaller than Q4KM, a little worse.Smaller footprint
Ling-3.0-tiny-IQ4_NL.gguf4.22 GBNon-linear 4-bit; good on hardware without fast K-quant kernels.Specialized hardware
Ling-3.0-tiny-IQ4_XS.gguf3.99 GBBest sub-4.5bpw option; usually beats Q4KS at a smaller size.Size-critical
Ling-3.0-tiny-Q4_1.gguf4.66 GBLegacy. Prefer Q4KM.Historical compatibility
Ling-3.0-tiny-Q4_0.gguf4.22 GBLegacy round-to-nearest. Prefer Q4KM unless a runtime needs this.Historical compatibility
Ling-3.0-tiny-MXFP4_MOE.gguf4.39 GBMoE-only 4-bit microscaling format for the expert tensors.Specialized MoE deployment
Ling-3.0-tiny-Q3_K_L.gguf3.86 GBSmall, with real quality loss. Usable when RAM is tight.Tight RAM constraints
Ling-3.0-tiny-Q3_K_M.gguf3.58 GBSmaller again; noticeable degradation.Memory-constrained
Ling-3.0-tiny-IQ3_M.gguf3.31 GBStrong at ~3.7bpw, clearly better than Q3KM.Quality-conscious sizing
Ling-3.0-tiny-IQ3_S.gguf3.27 GBSlightly smaller than IQ3_M.Compact version
Ling-3.0-tiny-Q3_K_S.gguf3.27 GBAggressive. Prefer IQ3_M at a similar size.Maximum compression
Ling-3.0-tiny-IQ3_XS.gguf3.11 GBAggressive but coherent.Extreme compression
Ling-3.0-tiny-IQ3_XXS.gguf2.91 GBVery aggressive; the last coherent step down.Extreme compression
Ling-3.0-tiny-Q2_K.gguf2.78 GBVery small, heavily degraded. For experimentation.Experimental only
Ling-3.0-tiny-IQ2_M.gguf2.52 GBThe smallest size most people find usable.Memory-constrained
Ling-3.0-tiny-Q2_K_S.gguf2.59 GBSmaller than Q2_K, at a further quality cost.Even smaller
Ling-3.0-tiny-IQ2_S.gguf2.31 GBBelow the usual usability line.Extreme compression
Ling-3.0-tiny-IQ2_XS.gguf2.27 GBExperimental.Experimental only
Ling-3.0-tiny-IQ2_XXS.gguf2.06 GBExperimental.Experimental only
Ling-3.0-tiny-Q2_0.gguf2.28 GBExtreme, group-64. Included for completeness.Historical compatibility
Ling-3.0-tiny-IQ1_M.gguf1.80 GBExtreme. Expect substantial degradation.Maximum compression
Ling-3.0-tiny-IQ1_S.gguf1.64 GBExtreme. Expect substantial degradation.Maximum compression
Ling-3.0-tiny-Q1_0.gguf1.21 GBExtreme. Included for completeness.Historical compatibility

All files are cut from the BF16 source (14.72 GB). Each quant is independently uploaded and deleted immediately after successful upload to minimize peak disk usage.

Model Architecture

Ling-3.0-tiny features a unique hybrid architecture combining KDA/MLA attention with a sparse MoE FFN:

Attention Mechanism

  • KDA (Kimi Delta Attention): 3 layers per 4-layer block
  • MLA (Multi-Head Latent Attention): 1 layer per 4-layer block
  • Hybrid Stacking: 3:1 ratio for efficient long-context processing

Feed-Forward Network

  • MoE (MultiplE Experts): 128 total experts
  • Routed Experts: 8 activated per token
  • Shared Expert: 1 additional expert
  • Activation Efficiency: Only 1.3B of 7.9B parameters activated per token

Core Components

  • Layers: 24 total
  • Hidden Size: 1536
  • Vocab Size: 157184
  • Position Embedding: Rotary Position Embedding (RoPE)
  • Precision: bfloat16 (source)

Inference

Generation Parameters

Important: Use temperature=1.0 and top_p=0.95 across all tasks and serving backends, including general chat, reasoning, and tool calling.
ParameterValueNotes
temperature1.0Required for all modes
top_p0.95Nucleus sampling threshold
top_k20Recommended for stable generation
max_new_tokens8192Thinking mode (increase for complex reasoning)
max_new_tokens2048Non-thinking mode
do_sampleTrueRequired when temperature > 0

Thinking Modes

ModeTemplate ParametersBehavior
Thinking (default)enable_thinking=TrueFull chain-of-thought reasoning inside <think>...</think>
Non-thinkingenable_thinking=FalseDirect answer with no reasoning overhead
Low-effortenable_thinking=True, low_effort=TrueBrief reasoning for simpler queries

Serving with llama.cpp

Basic Usage

bash
# Download a quantization (recommended: Q4_K_M)
huggingface-cli download NANI-Nithin/Ling-3.0-tiny-GGUF Ling-3.0-tiny-Q4_K_M.gguf --local-dir .

# Run inference
llama-cli -m Ling-3.0-tiny-Q4_K_M.gguf -p "Hello, how are you?"

Or use the Hugging Face Hub integration:

bash
# Serve with llama-server
llama-server -hf NANI-Nithin/Ling-3.0-tiny-GGUF:Q4_K_M

# Run with model path
llama-cli -hf NANI-Nithin/Ling-3.0-tiny-GGUF:Q4_K_M -p "Hello"

API Usage

bash
# Start the server
curl -s http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "granite-4.2-3b-Q4_K_M",
    "messages": [{"role": "user", "content": "Explain quantum computing in simple terms"}],
    "temperature": 1.0,
    "top_p": 0.95,
    "max_tokens": 8192
  }'

Quick Start Example

bash
# Download the model
huggingface-cli download NANI-Nithin/Ling-3.0-tiny-GGUF Ling-3.0-tiny-Q4_K_M.gguf

# Run inference
./llama-cli -m Ling-3.0-tiny-Q4_K_M.gguf -p "What is the Riemann hypothesis?"

Technical Details

  • Source Model: inclusionAI/Ling-3.0-tiny
  • Author: inclusionAI
  • License: MIT
  • Architecture: BailingMoeV3ForCausalLM (hybrid KDA/MLA + sparse MoE)
  • Parameters: 7.9B total, 1.3B activated per token
  • Experts: 128 total (8 routed + 1 shared per token)
  • Context Length: 131K tokens (natively supports 128K)
  • Created: August 10, 2026
  • Languages: Multiple languages supported

Usage Notes

  1. 1.Disk Space: Download one quantization at a time. Each file ranges from 1.21 GB (Q1_0) to 14.72 GB (BF16 source).
  2. 2.Memory Requirements: Varies by quantization. Q4KM requires ~4 GB VRAM, Q8_0 requires ~8 GB VRAM.
  3. 3.MoE Optimization: The sparse MoE architecture provides efficient inference while maintaining broad capabilities.
  4. 4.Thinking Mode: Enable enable_thinking=True to get chain-of-thought reasoning. Set to False for faster, direct answers.

Performance Characteristics

Ling-3.0-tiny achieves impressive efficiency:

  • FP8 Performance: ~100-105 tokens/s on DGX Spark, 86-90 tokens/s on M4 Pro MacBook
  • Memory Usage: ~8.34 GB peak at 8K context length
  • Agentic Performance: Score of 25 on Artificial Analysis Intelligence Index v4.1.1
  • End-to-End Latency: ~18 seconds for 500-token response including reasoning

Model Card Information

This GGUF repo contains quantized versions of Ling-3.0-tiny, featuring a unique hybrid architecture combining KDA/MLA attention with sparse MoE for efficient reasoning and agentic capabilities. The quantization was performed using llama.cpp's quantization pipeline, preserving the model's MoE efficiency while reducing size for deployment.

For the full source model documentation, including detailed training methodology, evaluation benchmarks, and advanced deployment recipes (SGLang, vLLM, Ollama), refer to the source repo: inclusionAI/Ling-3.0-tiny.


Note: These files are optimized for llama.cpp and are not compatible with vLLM, SGLang, or the Transformers library in their current format. For those frameworks, use the source model inclusionAI/Ling-3.0-tiny.