G3nadh/dgx-spark-benchmarks
DGX Spark LLM Benchmarks First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell). Hardware GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4) Memory: 128GB unified LPDDR5x (273 GB/s) CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725) Storage: 4TB NVMe Framework: Ollama 0.18.3 CUDA: 13.0 | Driver: 580.142 Benchmark Results Run 1 — General Inference (11 models) Model Size Prompt tok/s Gen tok/s Load Time Llama 3.1… See the full description on the dataset page: https://huggingface.co/datasets/G3nadh/dgx-spark-benchmarks.
168
DGX Spark LLM Benchmarks
First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell).
Hardware
- GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4)
- Memory: 128GB unified LPDDR5x (273 GB/s)
- CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725)
- Storage: 4TB NVMe
- Framework: Ollama 0.18.3
- CUDA: 13.0 | Driver: 580.142
Benchmark Results
Run 1 — General Inference (11 models)
Run 2 — Coding Benchmark
Run 3 — Context Scaling
Run 4 — Vision
Key Findings
- 27-32B is the sweet spot — 10-12 tok/s, genuinely interactive
- Prompt eval scales 3-4x with longer prompts on unified memory
- DeepSeek-R1 generates 10,710 tokens of reasoning for one coding question
- 90B vision model runs on a desktop at 3.47 tok/s
- 123B is the ceiling — Mistral Large at 2.28 tok/s barely interactive
- Generation speed is constant regardless of prompt length
Author
Gopi Trinadh Maddikunta
- University of Houston · MS Engineering Data Science
- Research Assistant, Dr. Peizhu Qian
- GSoC 2025 Contributor (Scala Center)
- GitHub: GOPITRINADH3561
- Website: gopitrinadh.site
