khyentsevision/tibetan-ai-leaderboard
Add a rebuildable lock file
Fix lint job: install git in python:3.12-slim before diffing
Require Google-style docstrings on public symbols via ruff
Add ruff lint gate scoped to changed files
Cut Translation Quality Metrics prose to a bare title line
Trim study methodology/findings prose pending peer review; add mitra-qwen35-2b-embedder metric
Translation Quality Metrics: strip LLM sub-tab prose, pills for judge rows
Add LLM-as-judge sub-tab: GEMBA-DA across twelve models
updated gitignore
Switch Tibetan Coverage to frequency-weighted scoring
Translation Quality Metrics: default sort by Pearson r, swap column order
Add 5 new reference-free metrics, Tibetan tokenizer coverage column
Rename MT Metrics tab to Translation Quality Metrics
MT Metrics: add ref-based/ref-free badges in All view, fix description
Add scripts/README.md: how to add metrics and models to MT Metrics tab
Add metric computation scripts from human-eval-data-study
Add All subtab to MT Metrics combining reference-based and reference-free
Refine MT Metrics tab: tooltips, remove category column, fix value display
Add MT Metrics tab and fix LanguageBench language pair description
Add GitHub link to DharmaBench tab
Add paper links to DharmaBench and TLUE tabs
Add LanguageBench description block and links to all tabs
Update header description and remove LanguageBench link
Add TNCC, WCM, MiLiC-Eval, TibetanQA tabs; add interpretive notes to all tabs
Add benchmark result files for TNCC, WCM, TibetanQA, TUSA/TNEC, MiLiC-Eval, Banzhida
Expand README with full benchmark descriptions and app architecture
Add Medal rankings and ScoreField bars to DharmaBench and TLUE tabs
Update About dialog to cover all three benchmark tabs
Move DharmaBench avg column left, expand benchmark descriptions
Add DharmaBench and TLUE benchmark tabs
Filter datasets and columns to Tibetan results only
Initial Tibetan-only leaderboard
initial commit
