datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dgx-spark-benchmarks
DGX Spark LLM Arena benchmarks
Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.dgx-spark-benchmarks
DGX Spark LLM Benchmarks
First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell).
Hardware
GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4)
Memory: 128GB unified LPDDR5x (273 GB/s)
CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725)
Storage: 4TB NVMe
Framework: Ollama 0.18.3
CUDA: 13.0 | Driver: 580.142
Benchmark Results
Run 1 — General Inference (11 models)
Model
Size
Prompt tok/s
Gen tok/s
Load Time
Llama 3.1 8B
4.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/G3nadh/dgx-spark-benchmarks.dgx-spark-moe-benchmarks
Four MoE models on a DGX Spark: speed, tool-calling, and what actually breaks
Full measurement campaign on NVIDIA DGX Spark (GB10, 128 GB unified, ~273 GB/s),
vLLM 0.23.1rc1.dev301+g04c2a8dea, arm64/sm121. Every number here is measured on this
hardware, with the raw evidence included.
The headline: on synthetic tool-calling benchmarks all four models score 91-95 %. In a
real coding agent, three of them score 0-1 out of 14 and one scores 11 out of 14.
If you pick a model from the… See the full description on the dataset page: https://huggingface.co/datasets/pocharlies/dgx-spark-moe-benchmarks.dgx-spark-eval
DGX Spark Model Evaluations
75 Messläufe in fünf Konfigurationen, alle auf einer Maschine gemessen.
Keine Herstellerangaben — jede Zahl stammt aus einem eigenen Lauf. Stand: 2026-08-17.
Die Website zu denselben Daten: https://results.southbyte.de/
Was gemessen wurde
Config
Zeilen
Inhalt
llm_local
20
Sprachmodelle, lokal mit vLLM serviert
llm_saas
28
dieselben Testfälle gegen Frontier-APIs, als Referenzrahmen
guardrails
5
Guard-Modelle gegen einen… See the full description on the dataset page: https://huggingface.co/datasets/SouthByte/dgx-spark-eval.nvidia-dgx-best-practicesdgxupdatedgxtestnemo-dgxchen-tong-cot-sft
Nemotron DGXChen/Tong CoT SFT Dataset
This repository packages the CoT training data used for the first
dgxchen-tong-unsloth-r32-2xrtxpro6000 SFT run that produced the 0.83
Kaggle adapter continuation point.
Provenance
Local source file: data/external/dgxchen_nemotron_cot_tong/problem_ids_matched.csv
Public upstream Kaggle dataset: dgxchen/nemotron-cot-tong
SFT config in the training repo: configs/sft/unsloth_dgxchen_2x_rtxpro6000.toml
Training adapter lineage:… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-dgxchen-tong-cot-sft.
