rvashurin/AmbigQA-clarifications-luna-filtered
AmbigQA Luna clarifications Corrected 25 September 2026: clarification generation is gated by Luna's judgement, not by the original dataset label. The dataset retains 1,971 questions, each with five clarification entries. 23 questions were judged sufficiently specified by Luna: their original question is repeated five times. 1,948 questions were judged underspecified by Luna: their lists contain generated clarifications. The is_underspecified column remains the original AmbigQA… See the full description on the dataset page: https://huggingface.co/datasets/rvashurin/AmbigQA-clarifications-luna-filtered.
AmbigQA Luna clarifications
Corrected 25 September 2026: clarification generation is gated by Luna's judgement, not by the original dataset label. The dataset retains 1,971 questions, each with five clarification entries.
- 23 questions were judged sufficiently specified by Luna: their original question is repeated five times.
- 1,948 questions were judged underspecified by Luna: their lists contain generated clarifications.
- The
is_underspecifiedcolumn remains the original AmbigQA ground-truth label, not Luna's prediction. It contains 1,141 true and 830 false values. A false source label can therefore have generated clarifications.
Generation policy
Use the model's first completed response saved in the accuracy evaluation. No clarification needed. selects original-question repetition. A nonempty clarification list selects generation. Do not override this judgement with the ground-truth label or later resamples.
The 811 source-negative questions that Luna called underspecified now use their saved accuracy-run clarification lists, followed by fresh responses only where necessary to reach five. The four source-positive questions that Luna called sufficiently specified now repeat the original question. The other saved generated lists are preserved.
Every request uses gpt-5.6-luna with the exact prompt.txt, substituting only {question}. Settings: Responses API, temperature 1.0, reasoning effort none, maximum output tokens 4,096. No original answers, labels, or earlier responses are supplied. Append responses in order until five clarification entries exist; keep the first five and retain duplicates. A later no-clarification-needed response supplies no entries and does not change the first decision.
Fields and retained rows
The same 31 questions previously excluded for selected yes/no forms or square brackets remain excluded. This correction preserves the 1,971 retained questions, their order, source labels and answers. It does not apply a new filter to the newly used accuracy-run outputs. correction_diagnostics.json records descriptive flags; compliance with every semantic prompt instruction is not guaranteed.
Classification accuracy
The saved first-response decisions are unchanged: 58.65% accuracy (1,156/1,971) against the original labels. TP=1,137, FN=4, FP=811, TN=19. Underspecified recall is 99.65%, specificity 2.29%, and balanced accuracy 50.97%. This measures the subset after the prior output-based filtering, not an unselected full benchmark.
classification_evaluation/accuracy_report.json.gz contains the original first-response decisions. Its metadata describes the evaluation before this correction. correction_audit.json.gz records the corrected per-row policy and supplemental generations; correction_verification.json verifies the current dataset. Parquet embeds correction and prior-evaluation provenance. removal_manifest.json and retained_index_map.json continue to identify the original excluded/retained source rows. Initial filter artifacts are historical and stored under provenance/.
Attribution
Derived from zykov/AmbigQA-clarifications, revision a80a7f4f2982f91e797238a97ef754ce3e2c1297, and AmbigQA, Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer (2020). Distributed under CC BY-SA 3.0. The supplied prompt is preserved verbatim; no claim about its original authorship is made here.
