Yobitel/google-gemma-2-9b-it__llm-inference-chatbot-short__019e3b973b4f
google/gemma-2-9b-it on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit TTFT P50 30.0542 ms TTFT P99 183.9538 ms TPOT P50 8.648 ms TPOT P99 10.7976 ms Total P50 Ms 1095.4794 Total P99 Ms 1422.5879 Req Per S Passing 3.3476 Req Per S All 3.3947 Compliance Rate 0.9861 Ok Rate 1 Throughput Tok Per S 385.0208 Power Avg W 901.0112 Power Peak W 937.794 Energy Joules… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/google-gemma-2-9b-it__llm-inference-chatbot-short__019e3b973b4f.
google/gemma-2-9b-it on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Headline metrics
Run configuration
- Model: google/gemma-2-9b-it @ unknown00
- Engine: vllm v0.21.0
- Quantization: fp16
- Hardware: NVIDIA H100 80GB HBM3
- Driver: 580.126.09
- CUDA: 13.0
- Run date: 2026-05-18T14:57:07.407582+00:00
- Seed: 42
Verification
This result is Sigstore-signed and Rekor-logged. Verify:
pip install inferencebench
bench verify hf://datasets/Yobitel/google-gemma-2-9b-it__llm-inference-chatbot-short__019e3b973b4f/envelope.jsonMethodology
See the suite methodology page.
Citation
@misc{inferencebench_019e3b973b4f,
title = { google/gemma-2-9b-it on llm.inference.chatbot-short },
author = { {InferenceBench community} },
year = { 2026 },
url = { https://huggingface.co/datasets/Yobitel/google-gemma-2-9b-it__llm-inference-chatbot-short__019e3b973b4f },
}Published via [InferenceBench](https://github.com/yobitelcomm/bench) — vendor-neutral AI benchmarks.
