datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dgx-spark-eval
DGX Spark Model Evaluations
75 Messläufe in fünf Konfigurationen, alle auf einer Maschine gemessen.
Keine Herstellerangaben — jede Zahl stammt aus einem eigenen Lauf. Stand: 2026-08-17.
Die Website zu denselben Daten: https://results.southbyte.de/
Was gemessen wurde
Config
Zeilen
Inhalt
llm_local
20
Sprachmodelle, lokal mit vLLM serviert
llm_saas
28
dieselben Testfälle gegen Frontier-APIs, als Referenzrahmen
guardrails
5
Guard-Modelle gegen einen… See the full description on the dataset page: https://huggingface.co/datasets/SouthByte/dgx-spark-eval.nova-industry-benchmark-results
Nova Industry Benchmark — Results
Model run outputs for the questions in
SparkSupernova/nova-industry-benchmark.
Why results live in their own repository
Results were previously stored as extra splits of the question dataset. Because each
model version wrote a different set of columns, adding the v5 run gave that dataset two
splits with incompatible schemas, and load_dataset failed for every consumer — including
the usage example on the model card.
Questions are a… See the full description on the dataset page: https://huggingface.co/datasets/SparkSupernova/nova-industry-benchmark-results.
