Yobitel/microsoft-phi-3-5-mini-instruct__llm-inference-chatbot-short__019e3b5cf170
microsoft/Phi-3.5-mini-instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit TTFT P50 18.4079 ms TTFT P99 265.3733 ms TPOT P50 4.6615 ms TPOT P99 6.5718 ms Total P50 Ms 531.3797 Total P99 Ms 786.9971 Req Per S Passing 6.5257 Req Per S All 6.6729 Compliance Rate 0.9779 Ok Rate 1 Throughput Tok Per S 716.0655 Power Avg W 861.866 Power Peak W 897.907 Energy… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/microsoft-phi-3-5-mini-instruct__llm-inference-chatbot-short__019e3b5cf170.
microsoft/Phi-3.5-mini-instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Headline metrics
Run configuration
- Model: microsoft/Phi-3.5-mini-instruct @ unknown00
- Engine: vllm v0.21.0
- Quantization: fp16
- Hardware: NVIDIA H100 80GB HBM3
- Driver: 580.126.09
- CUDA: 13.0
- Run date: 2026-05-18T13:53:27.408349+00:00
- Seed: 42
Verification
This result is Sigstore-signed and Rekor-logged. Verify:
pip install inferencebench
bench verify hf://datasets/Yobitel/microsoft-phi-3-5-mini-instruct__llm-inference-chatbot-short__019e3b5cf170/envelope.jsonMethodology
See the suite methodology page.
Citation
@misc{inferencebench_019e3b5cf170,
title = { microsoft/Phi-3.5-mini-instruct on llm.inference.chatbot-short },
author = { {InferenceBench community} },
year = { 2026 },
url = { https://huggingface.co/datasets/Yobitel/microsoft-phi-3-5-mini-instruct__llm-inference-chatbot-short__019e3b5cf170 },
}Published via [InferenceBench](https://github.com/yobitelcomm/bench) — vendor-neutral AI benchmarks.
