waltgrace/llama-cpp-expert-sniper
1
Q8 result: thrashes on 16 GB (CPU_REPACK doubles memory)
Add Gemma 4 benchmark results
Update MLX sniper link to 5.4 tok/s
Update MLX sniper link: 5.2 tok/s on 35B
Add full README with verified benchmarks and research findings
Add common/arg.cpp
Add common/common.h
Add common/common.cpp
Add src/CMakeLists.txt
Add src/llama-expert-cache.h
Add src/llama-expert-cache.cpp
Add src/llama-expert-cache-ctx.h
Add src/llama-expert-cache-ctx.cpp
initial commit
