madhavan02/llm-eval-harness
0
LLM Eval Harness — Model Comparison Dashboard
Interactive dashboard comparing llama-3.3-70b-versatile vs llama-3.1-8b-instant on Groq, evaluated on 10 AI/ML Q&A pairs.
Metrics shown
- P50 / P95 latency and time-to-first-token (TTFT)
- Tokens per second throughput
- Exact match and semantic similarity
- Hallucination rate and RAGAS faithfulness
- Cost per query and total cost
