lius-cc/Daoism-QA-Eval-v1
Daoism-QA-Eval-v1 Benchmark Evaluation Set for Daoism-Qwen3.5-9B and other LLMs on Daoist knowledge tasks. 由鼎稔道學館(lius.cc)發布。本 eval set 是 Daoism-QA-5K v0.1 中經 stratified sampling 抽出的 120 題 hold-out 集,永久切出不再用於任何 SFT 訓練。 完整評測方法論見本 repo 的 methodology.md 與 evaluator_prompt_v1.md。 規格 項目 值 樣本數 120 抽樣方式 Stratified(5 task_type × 24 題) 分層 每類依 groundedness_score 取 top/mid/bottom 1/3 各 8 題 來源 Daoism-QA-5K v0.1(249 條 pilot) 切出狀態 Hold-out,永久不再用於 SFT 訓練 語言… See the full description on the dataset page: https://huggingface.co/datasets/lius-cc/Daoism-QA-Eval-v1.
019
