CoolFace
Datasetpublic

axjns/strix-halo-inference-bench

Strix Halo Local Inference Benchmarks Measured prefill and decode throughput, and real VRAM cost, for local GGUF models on AMD Strix Halo (Radeon 8060S / gfx1151) under ROCm. Why this exists Strix Halo inverts the usual local-inference trade-off. A discrete 24 GB card gives you high memory bandwidth and a hard capacity ceiling; Strix Halo gives you the opposite — up to 64 GiB addressable as VRAM out of 128 GB unified, at substantially lower bandwidth. That changes… See the full description on the dataset page: https://huggingface.co/datasets/axjns/strix-halo-inference-bench.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
1likes36downloads
4 commits on main
ee68eb31mo ago

Lead with the dense-vs-MoE result; document unmeasurable models

axjns
9eecbce1mo ago

Clean sweep: 44 rows across 11 models

axjns
daccb291mo ago

Seed harness, dataset card, and validated first rows

axjns
e3242261mo ago

initial commit

axjns