gujilab/chinese-classical-bench
Chinese Classical Bench 中国古典语言能力评测基准 — 6 个任务 × 100 题 = 600 道,覆盖翻译、断句、字义、典故、续写填空、现代→文言压缩。 📊 在线排行榜: 🤗 Space — chinese-classical-bench-leaderboard 🔗 评测代码 & runner: github.com/gujilab/chinese-classical-bench — eval runner(OpenAI 兼容端点)、打分器、排行榜聚合脚本 📦 配套语料集: gujilab/chinese-classical-corpus (CC0 公有领域) — 题目均从该语料抽样生成 为什么做这个 中文(尤其文言文)常被说成"高密度优势"。这套基础设施(bench + corpus + 4 个论点实证实验)想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。 Tokenizer 层面(实证) 7 个主流 tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-bench.
data: psychometric audit — flag 18 circular char-gloss gold + 2 disputed (sync from bench v1.3)
data: psychometric audit — flag 18 circular char-gloss gold + 2 disputed (sync from bench v1.3)
data: psychometric audit — flag 18 circular char-gloss gold + 2 disputed (sync from bench v1.3)
README: replace 'free lunch' claim with empirical experiment results
docs: add tokenizer study cross-link + summary table
feat: add 6th task 'compress' — modern Chinese → classical compression
docs: reframe — Chinese classical as high-density language for LLM era
data: audit annotations (_audit_issue on 11 records)
data: audit annotations (_audit_issue on 11 records)
data: audit annotations (_audit_issue on 11 records)
data: audit annotations (_audit_issue on 11 records)
leaderboard: update findings narrative for DS/GLM-5 idiom tie
leaderboard: rescore idiom-source with smarter book matching (+17 hits across models, DS/GLM-5 now tied at 0.74)
chore: rename org dzxr/zi6me -> gujilab in links
docs: link to leaderboard Space
init: 5 tasks × 100 questions + dataset card with leaderboard
initial commit
