CoolFace
Datasetpublic

lius-cc/Daoism-QA-Eval-v1

Daoism-QA-Eval-v1 Benchmark Evaluation Set for Daoism-Qwen3.5-9B and other LLMs on Daoist knowledge tasks. 由鼎稔道學館(lius.cc)發布。本 eval set 是 Daoism-QA-5K v0.1 中經 stratified sampling 抽出的 120 題 hold-out 集,永久切出不再用於任何 SFT 訓練。 完整評測方法論見本 repo 的 methodology.md 與 evaluator_prompt_v1.md。 規格 項目 值 樣本數 120 抽樣方式 Stratified(5 task_type × 24 題) 分層 每類依 groundedness_score 取 top/mid/bottom 1/3 各 8 題 來源 Daoism-QA-5K v0.1(249 條 pilot) 切出狀態 Hold-out,永久不再用於 SFT 訓練 語言… See the full description on the dataset page: https://huggingface.co/datasets/lius-cc/Daoism-QA-Eval-v1.

sourceHugging Facecc-by-sa-4.0updated 4mo agoView on Hugging Face
0likes19downloads
5 commits on main
30f4c214mo ago

add eval set v1.0 (120 stratified samples)

chiyingliu
3ae42fc4mo ago

add evaluator_prompt_v1.md

chiyingliu
0fa6cb84mo ago

add methodology.md

chiyingliu
68e491e4mo ago

add README.md

chiyingliu
3a471254mo ago

initial commit

chiyingliu