CoolFace
Datasetpublic

G3nadh/dgx-spark-benchmarks

DGX Spark LLM Benchmarks First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell). Hardware GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4) Memory: 128GB unified LPDDR5x (273 GB/s) CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725) Storage: 4TB NVMe Framework: Ollama 0.18.3 CUDA: 13.0 | Driver: 580.142 Benchmark Results Run 1 — General Inference (11 models) Model Size Prompt tok/s Gen tok/s Load Time Llama 3.1… See the full description on the dataset page: https://huggingface.co/datasets/G3nadh/dgx-spark-benchmarks.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
1likes68downloads
Dataset Card

DGX Spark LLM Benchmarks

First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell).

Hardware

  • GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4)
  • Memory: 128GB unified LPDDR5x (273 GB/s)
  • CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725)
  • Storage: 4TB NVMe
  • Framework: Ollama 0.18.3
  • CUDA: 13.0 | Driver: 580.142

Benchmark Results

Run 1 — General Inference (11 models)

ModelSizePrompt tok/sGen tok/sLoad Time
Llama 3.1 8B4.9 GB574.7942.864.73s
Gemma3 27B17 GB164.6411.719.83s
Qwen2.5-Coder 32B19 GB288.9610.3615.13s
Qwen3 32B20 GB141.919.884.08s
CodeLlama 70B38 GB133.355.7328.81s
Nemotron 70B42 GB87.974.7727.25s
Llama 3.1 70B42 GB67.924.7628.35s
DeepSeek-R1 70B42 GB24.184.6847.02s
Llama 3.3 70B42 GB67.974.6627.41s
Qwen 2.5 72B47 GB122.194.4044.75s
Mistral Large 123B73 GB10.432.2886.14s

Run 2 — Coding Benchmark

ModelGen tok/sTokens Generated
Llama 3.1 8B42.51369
Qwen2.5-Coder 32B10.33552
Qwen3 32B9.388,180
CodeLlama 70B5.69752
Gemma3 27B11.61329
DeepSeek-R1 70B4.5110,710
Llama 3.3 70B4.67507

Run 3 — Context Scaling

ModelShort Prompt tok/sLong Prompt tok/sScaleGen tok/s
Llama 3.1 8B5742,0903.6x41.82
Gemma3 27B1646163.8x11.70
Qwen3 32B1414913.5x10.10
Llama 3.1 70B672123.2x4.69
Qwen 2.5 72B1222251.8x4.33
Nemotron 70B871641.9x4.62

Run 4 — Vision

ModelSizePrompt tok/sGen tok/sLoad Time
Llama3.2-Vision 90B54 GB6.023.4716.86s

Key Findings

  1. 1.27-32B is the sweet spot — 10-12 tok/s, genuinely interactive
  2. 2.Prompt eval scales 3-4x with longer prompts on unified memory
  3. 3.DeepSeek-R1 generates 10,710 tokens of reasoning for one coding question
  4. 4.90B vision model runs on a desktop at 3.47 tok/s
  5. 5.123B is the ceiling — Mistral Large at 2.28 tok/s barely interactive
  6. 6.Generation speed is constant regardless of prompt length

Author

Gopi Trinadh Maddikunta

  • University of Houston · MS Engineering Data Science
  • Research Assistant, Dr. Peizhu Qian
  • GSoC 2025 Contributor (Scala Center)
  • GitHub: GOPITRINADH3561
  • Website: gopitrinadh.site