xile42/lord-of-mysteries-fandom-evidence-sft
Lord of Mysteries Fandom Evidence SFT Dataset Overview This dataset provides evidence-aware training and retrieval material for building a Chinese Lord of the Mysteries knowledge assistant. The release is built from 425 Lord of the Mysteries Fandom Wiki pages. Source URLs, page titles, revision identifiers, and attribution metadata are preserved where available. The companion inference script can retrieve relevant source pages and attach exact source URLs before… See the full description on the dataset page: https://huggingface.co/datasets/xile42/lord-of-mysteries-fandom-evidence-sft.
Lord of Mysteries Fandom Evidence SFT Dataset
Overview
This dataset provides evidence-aware training and retrieval material for building a Chinese Lord of the Mysteries knowledge assistant.
The release is built from 425 Lord of the Mysteries Fandom Wiki pages. Source URLs, page titles, revision identifiers, and attribution metadata are preserved where available. The companion inference script can retrieve relevant source pages and attach exact source URLs before model generation.
This dataset is not built from raw novel chapters and must not be used to reconstruct, redistribute, or replace the original novel text.
Contents
The staged dataset repository includes:
train_sft.jsonlandvalidation_sft.jsonlfandom_lotm_pages.jsonlretrieval_aliases.jsoncurated_lotm_facts.jsondemo_questions.jsonlinference_example.pyrequirements-inference.txtsource_attribution_inventory.jsonsource_attribution_inventory.mdevaluation_reports/coverage_reports/hub_release_manifest.json
Data Sources
- Source type: Lord of the Mysteries Fandom Wiki pages
- Source count: 425 pages
- Primary language of sources: English
- Target assistant language: Chinese
- Source license: generally CC BY-SA 3.0 unless otherwise noted on the source page
- Source attribution: preserved in
source_attribution_inventory.jsonandsource_attribution_inventory.md
Relevant Fandom references:
- https://www.fandom.com/licensing
- https://community.fandom.com/wiki/Help:Licensing
Dataset Structure
Each SFT record uses a chat-style format:
{
"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
],
"source": {
"title": "...",
"url": "https://lordofthemysteries.fandom.com/wiki/...",
"revision_id": 123456,
"license": "CC BY-SA 3.0 unless otherwise noted on the source page"
}
}The retrieval files provide page-level source material and lightweight metadata for entity aliases, Chinese names, and disambiguation facts.
Splits
Evaluation and Coverage
The release includes evaluation and coverage reports for the recommended retrieval workflow:
These reports evaluate source coverage and retrieval behavior. They do not claim complete chapter-level coverage of the full novel series.
Usage
Install dependencies:
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
pip install huggingface_hubDownload the dataset:
hf download xile42/lord-of-mysteries-fandom-evidence-sft --repo-type dataset --local-dir lord-of-mysteries-fandom-evidence-sft
cd lord-of-mysteries-fandom-evidence-sft
pip install -r requirements-inference.txtRun retrieval-only checks:
python inference_example.py --retrieval-only --question "What is the source URL for Tarot Club?"
python inference_example.py --retrieval-only --questions-file demo_questions.jsonlRun generation with the paired LoRA adapter:
python inference_example.py \
--adapter xile42/qwen36-27b-lord-of-mysteries-lora \
--question "Introduce Tarot Club from Lord of the Mysteries and include the source."Recommended Pairing
This dataset is intended to be used with the paired model repository:
- Model:
xile42/qwen36-27b-lord-of-mysteries-lora - Base model:
Qwen/Qwen3.6-27B
The retrieval script defaults to:
fandom_lotm_pages.jsonlretrieval_aliases.jsoncurated_lotm_facts.jsonQwen/Qwen3.6-27B
Limitations
- The dataset is derived from wiki pages and may contain omissions or source-side inaccuracies.
- It does not represent official canon material.
- It does not include raw novel chapters.
- It should not be used to reproduce long copyrighted passages.
- Public redistribution should preserve attribution and follow the applicable CC BY-SA obligations.
License and Attribution
The dataset is released under CC BY-SA 3.0 for the derived Fandom-based material, subject to the terms and attribution requirements of the original source pages. Review the source attribution inventory before publishing or redistributing the dataset.
