cogeanu-marius/qwen3.6-27b-h100-bf16-benchmark
Two GPUs do not mean twice the users This study answers a serving decision, not a hardware trivia question: when a 27B model already fits on one H100, should a second GPU shard the model or run a second independent replica? The answer Concurrency is a load-generator setting, not a user count and not a promise. Capacity is the highest real arrival rate that satisfies a declared service objective. That is why this study measures both a saturated concurrency curve… See the full description on the dataset page: https://huggingface.co/datasets/cogeanu-marius/qwen3.6-27b-h100-bf16-benchmark.
This repository belongs to cogeanu-marius on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
