Jerrybro/bonsai2-27b-pq2-vs-fable-rtx5090
Bonsai 2 27B PQ2_0 vs Fable on a Single RTX 5090 This benchmark and reproducibility artifact compares single-GPU local LLM inference for Bonsai 2 27B / Qwen3.8 27B PQ2_0 GGUF through the Bonsai llama.cpp fork, its low-VRAM Q4_0 KV-cache quantization route with multimodal vision, and a Fable groupwise-int baseline. The controlled 210-question comparison on one RTX 5090 separates a quality-first Fable route from a 15,595 MiB sampled-peak Bonsai PQ2_0 + Q4_0-KV route.… See the full description on the dataset page: https://huggingface.co/datasets/Jerrybro/bonsai2-27b-pq2-vs-fable-rtx5090.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face