Santhiyarajan/omission-detection-synthetic
Omission Detection — Synthetic Sweep What is Omission Detection? Large language models (LLMs) in agentic pipelines often omit information present in their context window — they fail to surface a relevant fact even when it is theoretically visible. This dataset captures 75,876 controlled trials designed to measure and attribute these omissions across 9 taxonomic layers (L0–L8). Each trial generates a synthetic clinical document, embeds a "needle" fact at a… See the full description on the dataset page: https://huggingface.co/datasets/Santhiyarajan/omission-detection-synthetic.
010
