CoolFace
Datasetpublic

xile42/lord-of-mysteries-fandom-evidence-sft

Lord of Mysteries Fandom Evidence SFT Dataset Overview This dataset provides evidence-aware training and retrieval material for building a Chinese Lord of the Mysteries knowledge assistant. The release is built from 425 Lord of the Mysteries Fandom Wiki pages. Source URLs, page titles, revision identifiers, and attribution metadata are preserved where available. The companion inference script can retrieve relevant source pages and attach exact source URLs before… See the full description on the dataset page: https://huggingface.co/datasets/xile42/lord-of-mysteries-fandom-evidence-sft.

sourceHugging Facecc-by-sa-3.0updated 4mo agoView on Hugging Face
0likes69downloads
Dataset Card

Lord of Mysteries Fandom Evidence SFT Dataset

Overview

This dataset provides evidence-aware training and retrieval material for building a Chinese Lord of the Mysteries knowledge assistant.

The release is built from 425 Lord of the Mysteries Fandom Wiki pages. Source URLs, page titles, revision identifiers, and attribution metadata are preserved where available. The companion inference script can retrieve relevant source pages and attach exact source URLs before model generation.

This dataset is not built from raw novel chapters and must not be used to reconstruct, redistribute, or replace the original novel text.

Contents

The staged dataset repository includes:

  • —train_sft.jsonl and validation_sft.jsonl
  • —fandom_lotm_pages.jsonl
  • —retrieval_aliases.json
  • —curated_lotm_facts.json
  • —demo_questions.jsonl
  • —inference_example.py
  • —requirements-inference.txt
  • —source_attribution_inventory.json
  • —source_attribution_inventory.md
  • —evaluation_reports/
  • —coverage_reports/
  • —hub_release_manifest.json

Data Sources

  • —Source type: Lord of the Mysteries Fandom Wiki pages
  • —Source count: 425 pages
  • —Primary language of sources: English
  • —Target assistant language: Chinese
  • —Source license: generally CC BY-SA 3.0 unless otherwise noted on the source page
  • —Source attribution: preserved in source_attribution_inventory.json and source_attribution_inventory.md

Relevant Fandom references:

  • —https://www.fandom.com/licensing
  • —https://community.fandom.com/wiki/Help:Licensing

Dataset Structure

Each SFT record uses a chat-style format:

json
{
  "messages": [
    {"role": "system", "content": "..."},
    {"role": "user", "content": "..."},
    {"role": "assistant", "content": "..."}
  ],
  "source": {
    "title": "...",
    "url": "https://lordofthemysteries.fandom.com/wiki/...",
    "revision_id": 123456,
    "license": "CC BY-SA 3.0 unless otherwise noted on the source page"
  }
}

The retrieval files provide page-level source material and lightweight metadata for entity aliases, Chinese names, and disambiguation facts.

Splits

SplitExamples
train1246
validation108

Evaluation and Coverage

The release includes evaluation and coverage reports for the recommended retrieval workflow:

CheckResult
Fixed RAG evaluation12/12
Broad source retrieval evaluation49/49
Source page coverage audit425 pages
Broad key-anchor coverage49/49

These reports evaluate source coverage and retrieval behavior. They do not claim complete chapter-level coverage of the full novel series.

Usage

Install dependencies:

bash
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
pip install huggingface_hub

Download the dataset:

bash
hf download xile42/lord-of-mysteries-fandom-evidence-sft --repo-type dataset --local-dir lord-of-mysteries-fandom-evidence-sft
cd lord-of-mysteries-fandom-evidence-sft
pip install -r requirements-inference.txt

Run retrieval-only checks:

bash
python inference_example.py --retrieval-only --question "What is the source URL for Tarot Club?"
python inference_example.py --retrieval-only --questions-file demo_questions.jsonl

Run generation with the paired LoRA adapter:

bash
python inference_example.py \
  --adapter xile42/qwen36-27b-lord-of-mysteries-lora \
  --question "Introduce Tarot Club from Lord of the Mysteries and include the source."

Recommended Pairing

This dataset is intended to be used with the paired model repository:

  • —Model: xile42/qwen36-27b-lord-of-mysteries-lora
  • —Base model: Qwen/Qwen3.6-27B

The retrieval script defaults to:

  • —fandom_lotm_pages.jsonl
  • —retrieval_aliases.json
  • —curated_lotm_facts.json
  • —Qwen/Qwen3.6-27B

Limitations

  • —The dataset is derived from wiki pages and may contain omissions or source-side inaccuracies.
  • —It does not represent official canon material.
  • —It does not include raw novel chapters.
  • —It should not be used to reproduce long copyrighted passages.
  • —Public redistribution should preserve attribution and follow the applicable CC BY-SA obligations.

License and Attribution

The dataset is released under CC BY-SA 3.0 for the derived Fandom-based material, subject to the terms and attribution requirements of the original source pages. Review the source attribution inventory before publishing or redistributing the dataset.