strix-halo
DeepSeek-V4-Flash-Strix-Halo-GGUFQwen3.8-Flash-Next-MTP-Strix-Halo-GGUFqwen3.8-flash-next-gguf-strix-haloqwen3.8-27b-gguf-strix-haloMuse-Glimmer-30B-ROCmFP4-Strix-Halo-DFlash-GGUFDeepSeek-V4-Flash-Vision-Strix-Halo-GGUFdiffusiongemma-26b-a4b-it-strix-halo-fp16DeepSeek-V4-Flash-0731-StrixHalo-Verified-GGUF
strix-halo-inference-bench
Strix Halo Local Inference Benchmarks
Measured prefill and decode throughput, and real VRAM cost, for local GGUF models
on AMD Strix Halo (Radeon 8060S / gfx1151) under ROCm.
Why this exists
Strix Halo inverts the usual local-inference trade-off. A discrete 24 GB card gives
you high memory bandwidth and a hard capacity ceiling; Strix Halo gives you the
opposite — up to 64 GiB addressable as VRAM out of 128 GB unified, at substantially
lower bandwidth. That changes… See the full description on the dataset page: https://huggingface.co/datasets/axjns/strix-halo-inference-bench.strix-halo-bench-data
Strix Halo LLM Inference Benchmarks (gfx1151)
Reproducibility data for LLM inference benchmarks run on AMD Strix Halo hardware (Ryzen AI MAX+ 395 / Radeon 8060S iGPU / gfx1151, 128 GB unified memory). All logs are raw llama-bench and llama-server outputs from production benchmark runs documented in the strix-halo-llm-finetune-guide.
This dataset exists as a citation target — papers, blog posts, and forum threads referencing Strix Halo inference numbers can point to a stable… See the full description on the dataset page: https://huggingface.co/datasets/NorthstarAurora/strix-halo-bench-data.
