luoojason/mm-long-storytelling-bench
MM Long Storytelling Bench — v3 ⚠️ The 756 model-drafted questions have been WITHDRAWN from this dataset's splits (2026-08-05). They were drafted by a model that is also an evaluation target, which makes them circular as a measurement instrument. They are kept in full, with the reasoning, under data/v3/archive/ — nothing was deleted. The splits currently hold 6 worked examples (status: "example"), which document the required format and are not a benchmark. Do not use this… See the full description on the dataset page: https://huggingface.co/datasets/luoojason/mm-long-storytelling-bench.
v3 correction: withdraw the 756 model-drafted questions from the benchmark splits (kept in full under data/v3/archive/), leaving 6 worked examples. Questions are being rewritten by humans; the corpus and v2 are unchanged.
v3 correction: withdraw the 756 model-drafted questions from the benchmark splits (kept in full under data/v3/archive/), leaving 6 worked examples. Questions are being rewritten by humans; the corpus and v2 are unchanged. (part 3)
v3 correction: withdraw the 756 model-drafted questions from the benchmark splits (kept in full under data/v3/archive/), leaving 6 worked examples. Questions are being rewritten by humans; the corpus and v2 are unchanged. (part 2)
v3 correction: withdraw the 756 model-drafted questions from the benchmark splits (kept in full under data/v3/archive/), leaving 6 worked examples. Questions are being rewritten by humans; the corpus and v2 are unchanged.
v3: 125 sources (96 authors, 13.5M tokens), 756 questions across all 125 sources; multi-volume works merged into whole novels; 13 structural/grounding/leak checks passing. Draft, ungated, not human-verified. (part 2)
v3: 125 sources (96 authors, 13.5M tokens), 756 questions across all 125 sources; multi-volume works merged into whole novels; 13 structural/grounding/leak checks passing. Draft, ungated, not human-verified.
README: document v3 alongside v2
v3: 125 sources (96 authors, 13.5M tokens), 756 questions across all 125 sources; multi-volume works merged into whole novels; 13 structural/grounding/leak checks passing. Draft, ungated, not human-verified. (part 3)
v3: 125 sources (96 authors, 13.5M tokens), 756 questions across all 125 sources; multi-volume works merged into whole novels; 13 structural/grounding/leak checks passing. Draft, ungated, not human-verified. (part 2)
v3: 125 sources (96 authors, 13.5M tokens), 756 questions across all 125 sources; multi-volume works merged into whole novels; 13 structural/grounding/leak checks passing. Draft, ungated, not human-verified.
v2 update: 354 questions (118/type), 59 graded novels incl. baringgould (non-fiction, tagged); by_story refreshed; k1 document_ids
filler folder redesdale_further_memories: README notes intentionally no questions.jsonl
filler folder baringgould_freaks_fanaticism: README notes intentionally no questions.jsonl
remove 0-byte questions.jsonl from filler folder redesdale_further_memories (non-graded, no questions)
remove 0-byte questions.jsonl from filler folder baringgould_freaks_fanaticism (non-graded, no questions)
v2: add by_story view — one folder per novel with the questions whose gold it is (58 graded x 6, 2 filler)
v2: 60 obscure full-length PD novels, 348 aggregation questions (116/type, 6/novel); k1 document_ids + reconstruction recipe; raw sources; draft/ungated
v2: 60 obscure full-length PD novels, 348 aggregation questions (116/type, 6/novel), eval-ready k1 contexts + raw sources; draft/ungated
Add sourced/by_story — one folder per real story with its story text + questions
Point dataset card at sourced/ (real PD stories); drop deleted synthetic configs
rm synthetic sample
Delete generated synthetic stories (data/ pilot+v1) — superseded by sourced/
Add sourced/ — obscure public-domain stories + contamination-filter evidence + draft questions (replaces synthetic scaffolding)
Expand graded stories to real short narratives (~150-200w, RULES.md B6); re-gated clean
Add by_story view: one folder per story with its story text + the questions that use it
v1 rewrite: sibling-telling retrieval (no giveaways), per-story files, giveaway-clean gate
Fix config schema: hops always int, drop nullable decoys from eval rows
Normalize answer schema (answers:list<str> + answer_text) so configs load
MM Long Storytelling Bench: pilot + v1 contamination-safe core (draft/preview)
initial commit
