CoolFace
Apppublic

Raniahossam33/k2v3-error-scorecard

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes
App README

🧭 K2 V3 — Error-Analysis Scorecard

Per-example comparison of K2 V3 (4B / 9B) against size-matched reference models (Qwen2.5-3B / Qwen2.5-7B), evaluated under one validated lm-eval harness.

Each example is placed in one of four quadrants:

  • Both correct
  • K2 correct / reference wrong
  • Reference correct / K2 wrong
  • Both wrong

Benchmarks: HellaSwag, ARC-Challenge, Winogrande (loglikelihood) · GSM8K, BBH (CoT exact-match) · MATH (math-verify).

Features: overview KPIs, benchmark × quadrant heatmap, net per-example advantage chart, and a per-quadrant example browser.

Data is precomputed in scorecard.json by build_scorecard.py from the lm-eval per-example logs.