CoolFace
Datasetpublic

thoughtdag/context-repair-benchmark

ThoughtDAG Context Repair Benchmark What happens after one wrong assumption enters a long LLM conversation? This dataset turns context editing into a measurable intervention. Each synthetic case starts with a clean fact, introduces a false update, lets the error propagate through one to three downstream turns, and then asks the same final question under five graph conditions: clean polluted source_prune subgraph_prune recompute_descendants The central question is not only… See the full description on the dataset page: https://huggingface.co/datasets/thoughtdag/context-repair-benchmark.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes46downloads
5 commits on main
30ea4861mo ago

Polish dataset card

Chatchan
5408fb21mo ago

Add benchmark figures

Chatchan
7b781381mo ago

Add benchmark cases and reference results

Chatchan
5e0820b1mo ago

Add dataset card and manifest

Chatchan
7b13ac71mo ago

initial commit

Chatchan