CoolFace
Modelpublic

sarvamai/sarvam-30b-gguf

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
25likes454downloads
Model Card

image

!!! This is the GGUF version of Sarvam-30B !!!

Download the original weights here!

Index

  1. 1.Introduction
  2. 2.Architecture
  3. 3.Benchmarks
  4. 4.Knowledge & Coding
  5. 5.Reasoning & Math
  6. 6.Agentic
  7. 7.Inference
  8. 8.Hugging Face
  9. 9.vLLM
  10. 10.SGLang
  11. 11.Footnote
  12. 12.Citation

Introduction

Sarvam-30B is an advanced Mixture-of-Experts (MoE) model with 2.4B non-embedding active parameters, designed primarily for practical deployment. It combines strong reasoning, reliable coding ability, and best-in-class conversational quality across Indian languages. Sarvam-30B is built to run reliably in resource-constrained environments and can handle multilingual voice calls while performing tool calls.

A major focus during training was the Indian context and languages, resulting in state-of-the-art performance across 22 Indian languages for its model size.

Sarvam-30B is open-sourced under the Apache License. For more details, see our blog.

Architecture

The 30B MoE model is designed for throughput and memory efficiency. It uses 19 layers, a dense FFN intermediate_size of 8192, moe_intermediate_size of 1024, top-6 routing, grouped KV heads (num_key_value_heads=4), and an extremely high rope_theta (8e6) for long-context stability without RoPE scaling. It has 128 experts with a shared expert, a routed scaling factor of 2.5, and auxiliary-loss-free router balancing. The 30B model focuses on throughput and memory efficiency through fewer layers, grouped KV attention, and smaller experts.

Benchmarks

<details> <summary>Knowledge & Coding</summary>

BenchmarkSarvam-30BGemma 27B ItMistral-3.2-24BOLMo 3.1 32B ThinkNemotron-3-Nano-30B-A3BQwen3-30B-Thinking-2507GLM 4.7 FlashGPT-OSS-20B
Math50097.087.469.496.298.097.697.094.2
HumanEval92.188.492.995.197.695.796.395.7
MBPP92.781.878.358.791.994.391.895.3
Live Code Bench v670.028.026.073.068.366.064.061.0
MMLU85.181.280.586.484.088.486.985.3
MMLU Pro80.068.169.172.078.380.973.675.0
MILU76.869.267.969.964.882.675.673.7
Arena Hard v249.050.143.142.067.772.158.162.9
Writing Bench78.771.470.375.783.785.079.279.1

</details>

<details> <summary>Reasoning & Math</summary>

BenchmarkSarvam-30BOLMo 3.1 32BNemotron-3-Nano-30BQwen3-30B-Thinking-2507GLM 4.7 FlashGPT-OSS-20B
GPQA Diamond66.557.573.073.475.271.5
AIME 25 (w/ Tools)88.3 (96.7)78.1 (81.7)89.1 (99.2)85.0 (-)91.6 (-)91.7 (98.7)
HMMT (Feb 25)73.351.785.071.485.076.7
HMMT (Nov 25)74.258.375.073.381.768.3
Beyond AIME58.348.564.061.060.046.0

</details>

<details> <summary>Agentic</summary>

BenchmarkSarvam-30BNemotron-3-Nano-30BQwen3-30B-Thinking-2507GLM 4.7 FlashGPT-OSS-20B
BrowseComp35.523.82.942.828.3
SWE Bench Verified34.038.822.059.234.0
τ² Bench (avg.)45.749.047.779.548.7
See footnote for evaluation details.

</details>

Inference

Clone and build

bash
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp && cmake -B build && cmake --build build --config Release -j

Download the model (all shards)

bash
huggingface-cli download sarvamai/sarvam-30b-gguf --local-dir sarvam-30b-gguf

Run interactive chat

bash
./build/bin/llama-cli \
  -m sarvam-30b-gguf/sarvam-30b-Q4_K_M.gguf-00001-of-00006.gguf \
  -c 4096 \
  -n 512 \
  -p "You are a helpful assistant." \
  --conversation

OpenAI-compatible API server

bash
./build/bin/llama-server \
  -m sarvam-30b-gguf/sarvam-30b-Q4_K_M.gguf-00001-of-00006.gguf \
  -c 4096 \
  --host 0.0.0.0 \
  --port 8080

Then query it:

bash
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [
      {"role": "user", "content": "Explain quantum computing in simple terms."}
    ],
    "temperature": 0.8,
    "max_tokens": 512
  }'

Footnote

  • —General settings: All benchmarks are evaluated with a maximum context length of 65,536 tokens.
  • —Reasoning & Math benchmarks (Math500, MMLU, MMLU Pro, GPQA Diamond, AIME 25, Beyond AIME, HMMT, HumanEval, MBPP): Evaluated with temperature=1.0, top_p=1.0, max_new_tokens=65536.
  • —Coding & Knowledge benchmarks (Live Code Bench v6, Arena Hard v2, IF Eval): Evaluated with temperature=1.0, top_p=1.0, max_new_tokens=65536.
  • —Writing Bench: Responses generated using official Writing-Bench parameters: temperature=0.7, top_p=0.8, top_k=20, max_length=16000. Scoring performed using the official Writing-Bench critic model with: temperature=1.0, top_p=0.95, max_length=2048.
  • —Agentic benchmarks (BrowseComp, SWE Bench Verified, τ² Bench): Evaluated with temperature=0.5, top_p=1.0, max_new_tokens=32768.

Citation

@misc{sarvam_sovereign_models,
  title        = {Introducing Sarvam's Sovereign Models},
  author       = {{Sarvam Foundation Models Team}},
  year         = {2026},
  howpublished = {\url{https://www.sarvam.ai/blogs/sarvam-30b-105b}},
  note         = {Accessed: 2026-03-03}
}