CoolFace
Apppublic

madhavan02/llm-eval-harness

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes
App README

LLM Eval Harness — Model Comparison Dashboard

Interactive dashboard comparing llama-3.3-70b-versatile vs llama-3.1-8b-instant on Groq, evaluated on 10 AI/ML Q&A pairs.

Metrics shown

  • P50 / P95 latency and time-to-first-token (TTFT)
  • Tokens per second throughput
  • Exact match and semantic similarity
  • Hallucination rate and RAGAS faithfulness
  • Cost per query and total cost

Source

github.com/madhavanbalaji02/llm-eval-harness