short-inference
deepseek-ai-deepseek-coder-v2-lite-instruct__llm-inference-chatbot-short__019e3b6f22ca
deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
74.4268
ms
TTFT P99
521.1922
ms
TPOT P50
21.0371
ms
TPOT P99
23.458
ms
Total P50 Ms
2661.4077
Total P99 Ms
3136.5972
Req Per S Passing
1.0172
Req Per S All
1.1559
Compliance Rate
0.88
Ok Rate
1
Throughput Tok Per S
134.316
Power Avg W
808.3085
Power Peak W
854.62… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/deepseek-ai-deepseek-coder-v2-lite-instruct__llm-inference-chatbot-short__019e3b6f22ca.microsoft-phi-3-5-mini-instruct__llm-inference-chatbot-short__019e3b5cf170
microsoft/Phi-3.5-mini-instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
18.4079
ms
TTFT P99
265.3733
ms
TPOT P50
4.6615
ms
TPOT P99
6.5718
ms
Total P50 Ms
531.3797
Total P99 Ms
786.9971
Req Per S Passing
6.5257
Req Per S All
6.6729
Compliance Rate
0.9779
Ok Rate
1
Throughput Tok Per S
716.0655
Power Avg W
861.866
Power Peak W
897.907
Energy Joules… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/microsoft-phi-3-5-mini-instruct__llm-inference-chatbot-short__019e3b5cf170.google-gemma-2-9b-it__llm-inference-chatbot-short__019e3b973b4f
google/gemma-2-9b-it on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
30.0542
ms
TTFT P99
183.9538
ms
TPOT P50
8.648
ms
TPOT P99
10.7976
ms
Total P50 Ms
1095.4794
Total P99 Ms
1422.5879
Req Per S Passing
3.3476
Req Per S All
3.3947
Compliance Rate
0.9861
Ok Rate
1
Throughput Tok Per S
385.0208
Power Avg W
901.0112
Power Peak W
937.794
Energy Joules… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/google-gemma-2-9b-it__llm-inference-chatbot-short__019e3b973b4f.qwen-qwen2-vl-7b-instruct__llm-inference-chatbot-short__019e3b85e34e
Qwen/Qwen2-VL-7B-Instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
42.4998
ms
TTFT P99
198.2181
ms
TPOT P50
15.3479
ms
TPOT P99
15.7973
ms
Total P50 Ms
1981.2403
Total P99 Ms
2016.4637
Req Per S Passing
1.9266
Req Per S All
1.9724
Compliance Rate
0.9767
Ok Rate
1
Throughput Tok Per S
215.6387
Power Avg W
836.8075
Power Peak W
857.025
Energy Joules… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/qwen-qwen2-vl-7b-instruct__llm-inference-chatbot-short__019e3b85e34e.qwen-qwen2-5-7b-instruct__llm-inference-chatbot-short__019e3b395d04
Qwen/Qwen2.5-7B-Instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
26.8358
ms
TTFT P99
52.9781
ms
TPOT P50
6.2672
ms
TPOT P99
6.504
ms
Total P50 Ms
822.2015
Total P99 Ms
929.0652
Req Per S Passing
4.585
Req Per S All
4.6332
Compliance Rate
0.9896
Ok Rate
1
Throughput Tok Per S
540.8831
Power Avg W
908.6194
Power Peak W
947.635
Energy Joules… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/qwen-qwen2-5-7b-instruct__llm-inference-chatbot-short__019e3b395d04.meta-llama-llama-3-1-8b-instruct__llm-inference-chatbot-short__019e3b2fef77
meta-llama/Llama-3.1-8B-Instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
25.3868
ms
TTFT P99
1677.528
ms
TPOT P50
6.5951
ms
TPOT P99
10.8911
ms
Total P50 Ms
862.5176
Total P99 Ms
3020.5284
Req Per S Passing
1.8011
Req Per S All
1.8497
Compliance Rate
0.9737
Ok Rate
1
Throughput Tok Per S
227.5172
Cost Usd Per Million Tokens
0.0575
Cost Source… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-8b-instruct__llm-inference-chatbot-short__019e3b2fef77.
