CoolFace
Datasetpublic

cogeanu-marius/qwen3.6-27b-h100-bf16-benchmark

Two GPUs do not mean twice the users This study answers a serving decision, not a hardware trivia question: when a 27B model already fits on one H100, should a second GPU shard the model or run a second independent replica? The answer Concurrency is a load-generator setting, not a user count and not a promise. Capacity is the highest real arrival rate that satisfies a declared service objective. That is why this study measures both a saturated concurrency curve… See the full description on the dataset page: https://huggingface.co/datasets/cogeanu-marius/qwen3.6-27b-h100-bf16-benchmark.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes20downloads
settings

This repository belongs to cogeanu-marius on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameqwen3.6-27b-h100-bf16-benchmark
visibilitypublic
licenceapache-2.0
gatedno
ownercogeanu-marius
Account settings
cogeanu-marius/qwen3.6-27b-h100-bf16-benchmark · CoolFace