datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Daoism-QA-Eval-v1
Daoism-QA-Eval-v1
Benchmark Evaluation Set for Daoism-Qwen3.5-9B and other LLMs on Daoist knowledge tasks.
由鼎稔道學館(lius.cc)發布。本 eval set 是 Daoism-QA-5K v0.1 中經 stratified sampling 抽出的 120 題 hold-out 集,永久切出不再用於任何 SFT 訓練。
完整評測方法論見本 repo 的 methodology.md 與 evaluator_prompt_v1.md。
規格
項目
值
樣本數
120
抽樣方式
Stratified(5 task_type × 24 題)
分層
每類依 groundedness_score 取 top/mid/bottom 1/3 各 8 題
來源
Daoism-QA-5K v0.1(249 條 pilot)
切出狀態
Hold-out,永久不再用於 SFT 訓練
語言
繁體中文… See the full description on the dataset page: https://huggingface.co/datasets/lius-cc/Daoism-QA-Eval-v1.Daoism-Wiki-QA-280K
Daoism-Wiki-QA-280K
規模型道教問答資料集——263,910 條從鼎稔道學館 wiki 與論文庫批量蒸餾、經 critic_score ≥ 93 嚴格篩選後的高品質 QA pairs。
由鼎稔道學館(lius.cc)發布。本資料集與 Daoism-QA-5K(grounded 精選子集)形成「規模 × 精度」雙資產,配合 Daoism-Qwen3.5-9B 開源道教 LLM 完成「模型 + 資料」全棧開源。
規格
項目
值
樣本數
263,910
篩選門檻
critic_score >= 93
來源規模
578,174 raw → 263,910 final(保留率 45.6%)
平均 critic_score
~95.8
滿分 (100) 樣本
41,830(15.9%)
95-99 分樣本
174,006(65.9%)
93-94 分樣本
48,074(18.2%)
語言
繁體中文(zh-Hant)
檔案大小
240 MB
重複去除
4,769 條(依… See the full description on the dataset page: https://huggingface.co/datasets/lius-cc/Daoism-Wiki-QA-280K.CondAmbigQA
Dataset Card for CondAmbigQA
📢 Note: An expanded version with 2000 entries is now available at CondAmbigQA-2K
Dataset Description
CondAmbigQA is a specialized benchmark dataset containing 200 ambiguous queries with condition-aware evaluation metrics. It introduces "conditions" - contextual constraints that resolve ambiguities in question-answering tasks.
Supported Tasks
The dataset supports conditional question answering where systems must:
Identify… See the full description on the dataset page: https://huggingface.co/datasets/Apocalypse-AGI-DAO/CondAmbigQA.Daoism-QA-5K
Daoism-QA-5K
首個結構化、帶證據引用的道教問答資料集(v0.1 pilot preview — 249 條;v1.0 計畫 5,000 條)
由鼎稔道學館(lius.cc)發布,配合全球首個開源道教 LLM Daoism-Qwen3.5-9B 同步建設道教 NLP 標準語料庫。
概要
每條樣本是一組「問題 + 結構化回答 + 引用證據鏈」,所有可驗證主張都對應到 retrieval 來源的 evidence_id。資料來自鼎稔道學館內部 wiki(117,830 條條目、102,303 條 embedding)以及學術論文庫。
生成模型:OpenAI gpt-5.5(透過 codex-oauth proxy)
Retrieval:PostgreSQL FTS(NodeSearch.haystack gin_trgm_ops 索引)
驗證:Pydantic v2 schema + 6 個品質指標 hard validators
語言:繁體中文(zh-Hant)
規格
項目
值… See the full description on the dataset page: https://huggingface.co/datasets/lius-cc/Daoism-QA-5K.
