CoolFace
Datasetpublic

Yobitel/meta-llama-llama-3-1-8b-instruct__llm-inference-chatbot-short__019e3b2fef77

meta-llama/Llama-3.1-8B-Instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit TTFT P50 25.3868 ms TTFT P99 1677.528 ms TPOT P50 6.5951 ms TPOT P99 10.8911 ms Total P50 Ms 862.5176 Total P99 Ms 3020.5284 Req Per S Passing 1.8011 Req Per S All 1.8497 Compliance Rate 0.9737 Ok Rate 1 Throughput Tok Per S 227.5172 Cost Usd Per Million Tokens 0.0575 Cost Source… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-8b-instruct__llm-inference-chatbot-short__019e3b2fef77.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes10downloads
Dataset Card

meta-llama/Llama-3.1-8B-Instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)

Back to leaderboard

Headline metrics

MetricValueUnit
TTFT P5025.3868ms
TTFT P991677.528ms
TPOT P506.5951ms
TPOT P9910.8911ms
Total P50 Ms862.5176
Total P99 Ms3020.5284
Req Per S Passing1.8011
Req Per S All1.8497
Compliance Rate0.9737
Ok Rate1
Throughput Tok Per S227.5172
Cost Usd Per Million Tokens0.0575
Cost Sourceregistry:groq
Power Avg W758.3006
Power Peak W943.221
Energy Joules Total15249.7971
Joules per token3.2627J
Slo Hardware Classh100
Slo Template Resolvedttft<200ms, tpot<50ms, total<3000ms

Run configuration

  • —Model: meta-llama/Llama-3.1-8B-Instruct @ unknown00
  • —Engine: vllm v0.21.0
  • —Quantization: fp16
  • —Hardware: NVIDIA H100 80GB HBM3
  • —Driver: 580.126.09
  • —CUDA: 13.0
  • —Run date: 2026-05-18T13:04:17.783238+00:00
  • —Seed: 42

Verification

This result is Sigstore-signed and Rekor-logged. Verify:

bash
pip install inferencebench
bench verify hf://datasets/Yobitel/meta-llama-llama-3-1-8b-instruct__llm-inference-chatbot-short__019e3b2fef77/envelope.json

Rekor entry: log index -1

Methodology

See the suite methodology page.

Citation

bibtex
@misc{inferencebench_019e3b2fef77,
  title = { meta-llama/Llama-3.1-8B-Instruct on llm.inference.chatbot-short },
  author = { {InferenceBench community} },
  year = { 2026 },
  url = { https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-8b-instruct__llm-inference-chatbot-short__019e3b2fef77 },
}

Published via [InferenceBench](https://github.com/yobitelcomm/bench) — vendor-neutral AI benchmarks.