CoolFace
17 results

DGX-Spark

Djangodevreng /dgx-spark-benchmarks DGX Spark LLM Arena benchmarks Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.tabularn<1K1 likes133 downloads6d agoHugging Facebonellisystems /dgx-spark-nvfp4-notes dgx-spark-nvfp4-notes Assets for the Hugging Face discussion on nvidia/Qwen3.6-35B-A3B-NVFP4: official DGX Spark Marlin recipe dies under concurrent json_schema load on GB10 / SM121. imagen<1K0 likes98 downloads1mo agoHugging FaceG3nadh /dgx-spark-benchmarks DGX Spark LLM Benchmarks First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell). Hardware GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4) Memory: 128GB unified LPDDR5x (273 GB/s) CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725) Storage: 4TB NVMe Framework: Ollama 0.18.3 CUDA: 13.0 | Driver: 580.142 Benchmark Results Run 1 — General Inference (11 models) Model Size Prompt tok/s Gen tok/s Load Time Llama 3.1 8B 4.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/G3nadh/dgx-spark-benchmarks.texttext-generationn<1K1 likes71 downloads6mo agoHugging FaceH-K-B /glm-5.3-flash-dgx-spark-vllm GLM-5.3-Flash on 2x NVIDIA DGX Spark (GB10): a working multi-node vLLM config TL;DR: GLM-5.3-Flash (NVFP4, modelopt quant) running tensor-parallel across two DGX Spark GB10 nodes over 200GbE RoCEv2 RDMA, with NVFP4 KV-cache, CUDA graphs, and MTP-3 speculative decoding: ~20-23 tok/s single-stream, ~49 tok/s aggregate at 4 concurrent requests on 64K context. About 2x over a naive TCP/eager launch. This repo documents the exact configuration, the four vLLM patches it needs… See the full description on the dataset page: https://huggingface.co/datasets/H-K-B/glm-5.3-flash-dgx-spark-vllm.0 likes71 downloads8d agoHugging FaceNathan-Maine /dgx-spark-kv-cache-benchmark KV Cache Quantization on NVIDIA DGX Spark GB10 Corrected benchmarks (v3, April 2026) — KV cache quantization behavior on the NVIDIA DGX Spark's GB10 Grace Blackwell unified memory architecture. Author: Nathan Maine Date: March 2026, corrected April 2026 Hardware: NVIDIA DGX Spark (GB10, compute 12.1, 128GB unified memory) Correction Notice: The original v1 benchmarks (March 31) contained methodology errors. Memory was measured via RSS (wrong on unified memory) and some… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/dgx-spark-kv-cache-benchmark.othern<1K3 likes62 downloads4mo agoHugging Facepocharlies /dgx-spark-moe-benchmarks Four MoE models on a DGX Spark: speed, tool-calling, and what actually breaks Full measurement campaign on NVIDIA DGX Spark (GB10, 128 GB unified, ~273 GB/s), vLLM 0.23.1rc1.dev301+g04c2a8dea, arm64/sm121. Every number here is measured on this hardware, with the raw evidence included. The headline: on synthetic tool-calling benchmarks all four models score 91-95 %. In a real coding agent, three of them score 0-1 out of 14 and one scores 11 out of 14. If you pick a model from the… See the full description on the dataset page: https://huggingface.co/datasets/pocharlies/dgx-spark-moe-benchmarks.textn<1K0 likes59 downloads2mo agoHugging Face