CoolFace
Datasetpublic

luoojason/mm-long-storytelling-bench

MM Long Storytelling Bench — v3 ⚠️ The 756 model-drafted questions have been WITHDRAWN from this dataset's splits (2026-08-05). They were drafted by a model that is also an evaluation target, which makes them circular as a measurement instrument. They are kept in full, with the reasoning, under data/v3/archive/ — nothing was deleted. The splits currently hold 6 worked examples (status: "example"), which document the required format and are not a benchmark. Do not use this… See the full description on the dataset page: https://huggingface.co/datasets/luoojason/mm-long-storytelling-bench.

sourceHugging Facecc0-1.0updated 2mo agoView on Hugging Face
0likes216downloads
30 commits on main
95d098e2mo ago

v3 correction: withdraw the 756 model-drafted questions from the benchmark splits (kept in full under data/v3/archive/), leaving 6 worked examples. Questions are being rewritten by humans; the corpus and v2 are unchanged.

luoojason
4e0b7422mo ago

v3 correction: withdraw the 756 model-drafted questions from the benchmark splits (kept in full under data/v3/archive/), leaving 6 worked examples. Questions are being rewritten by humans; the corpus and v2 are unchanged. (part 3)

luoojason
27ce3f72mo ago

v3 correction: withdraw the 756 model-drafted questions from the benchmark splits (kept in full under data/v3/archive/), leaving 6 worked examples. Questions are being rewritten by humans; the corpus and v2 are unchanged. (part 2)

luoojason
308af202mo ago

v3 correction: withdraw the 756 model-drafted questions from the benchmark splits (kept in full under data/v3/archive/), leaving 6 worked examples. Questions are being rewritten by humans; the corpus and v2 are unchanged.

luoojason
a91c0212mo ago

v3: 125 sources (96 authors, 13.5M tokens), 756 questions across all 125 sources; multi-volume works merged into whole novels; 13 structural/grounding/leak checks passing. Draft, ungated, not human-verified. (part 2)

luoojason
86ddb3c2mo ago

v3: 125 sources (96 authors, 13.5M tokens), 756 questions across all 125 sources; multi-volume works merged into whole novels; 13 structural/grounding/leak checks passing. Draft, ungated, not human-verified.

luoojason
cfba3952mo ago

README: document v3 alongside v2

luoojason
885937a2mo ago

v3: 125 sources (96 authors, 13.5M tokens), 756 questions across all 125 sources; multi-volume works merged into whole novels; 13 structural/grounding/leak checks passing. Draft, ungated, not human-verified. (part 3)

luoojason
1244d882mo ago

v3: 125 sources (96 authors, 13.5M tokens), 756 questions across all 125 sources; multi-volume works merged into whole novels; 13 structural/grounding/leak checks passing. Draft, ungated, not human-verified. (part 2)

luoojason
66510ca2mo ago

v3: 125 sources (96 authors, 13.5M tokens), 756 questions across all 125 sources; multi-volume works merged into whole novels; 13 structural/grounding/leak checks passing. Draft, ungated, not human-verified.

luoojason
02e892b2mo ago

v2 update: 354 questions (118/type), 59 graded novels incl. baringgould (non-fiction, tagged); by_story refreshed; k1 document_ids

luoojason
b8216512mo ago

filler folder redesdale_further_memories: README notes intentionally no questions.jsonl

luoojason
99b9c982mo ago

filler folder baringgould_freaks_fanaticism: README notes intentionally no questions.jsonl

luoojason
dbcbfdd2mo ago

remove 0-byte questions.jsonl from filler folder redesdale_further_memories (non-graded, no questions)

luoojason
c5ae65b2mo ago

remove 0-byte questions.jsonl from filler folder baringgould_freaks_fanaticism (non-graded, no questions)

luoojason
fb7900f2mo ago

v2: add by_story view — one folder per novel with the questions whose gold it is (58 graded x 6, 2 filler)

luoojason
9f9bce02mo ago

v2: 60 obscure full-length PD novels, 348 aggregation questions (116/type, 6/novel); k1 document_ids + reconstruction recipe; raw sources; draft/ungated

luoojason
bea9a082mo ago

v2: 60 obscure full-length PD novels, 348 aggregation questions (116/type, 6/novel), eval-ready k1 contexts + raw sources; draft/ungated

luoojason
356b3be3mo ago

Add sourced/by_story — one folder per real story with its story text + questions

luoojason
e37f0603mo ago

Point dataset card at sourced/ (real PD stories); drop deleted synthetic configs

luoojason
b2c801a3mo ago

rm synthetic sample

luoojason
0ee36f23mo ago

Delete generated synthetic stories (data/ pilot+v1) — superseded by sourced/

luoojason
bab4bfe3mo ago

Add sourced/ — obscure public-domain stories + contamination-filter evidence + draft questions (replaces synthetic scaffolding)

luoojason
37bdc6c3mo ago

Expand graded stories to real short narratives (~150-200w, RULES.md B6); re-gated clean

luoojason
ed1f3433mo ago

Add by_story view: one folder per story with its story text + the questions that use it

luoojason
95e060c3mo ago

v1 rewrite: sibling-telling retrieval (no giveaways), per-story files, giveaway-clean gate

luoojason
6e946ef3mo ago

Fix config schema: hops always int, drop nullable decoys from eval rows

luoojason
b0a5cad3mo ago

Normalize answer schema (answers:list<str> + answer_text) so configs load

luoojason
f4a12ad3mo ago

MM Long Storytelling Bench: pilot + v1 contamination-safe core (draft/preview)

luoojason
4df3f6e3mo ago

initial commit

luoojason