datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coqar-clarifications-audio
CoQAR Clarifications with synthetic context audio
These audio recordings are AI-generated speech, not recordings of human speakers.
OpenAI tts-1 narrated each exact story using voice alloy, speed 1,
and MP3 output. Long stories are synthesized in ordered parts and joined; see the
audio generation manifest for part boundaries and measured audio properties.
No questions, answers, rationales, or stored model prompts were narrated.
The original appended and inserted configurations… See the full description on the dataset page: https://huggingface.co/datasets/rvashurin/coqar-clarifications-audio.coqar-clarifications
CoQAR Clarifications
This dataset pairs 1,000 CoQAR development questions with their original stories and stories damaged by sentence deletion. Each of the resulting 2,000 inputs has five sampled model clarifications. Two configurations reuse the same generated additions and differ only in where those additions are placed.
Configuration
Rows in dev
Clarifications per row
Placement
appended
2,000
5
At the end of the input story
inserted
2,000
5
At the deleted passage… See the full description on the dataset page: https://huggingface.co/datasets/rvashurin/coqar-clarifications.AmbigQA-clarifications-luna-filtered
AmbigQA Luna clarifications
Corrected 25 September 2026: clarification generation is gated by Luna's judgement, not by the original dataset label. The dataset retains 1,971 questions, each with five clarification entries.
23 questions were judged sufficiently specified by Luna: their original question is repeated five times.
1,948 questions were judged underspecified by Luna: their lists contain generated clarifications.
The is_underspecified column remains the original AmbigQA… See the full description on the dataset page: https://huggingface.co/datasets/rvashurin/AmbigQA-clarifications-luna-filtered.AmbigQA-clarifications
AmbigQA clarifications
The dataset contains 2,002 questions from the full configuration, validation split of AmbigQA.
Field
Description
question
Original question from AmbigQA, unchanged.
clarifications
A list of five clarifications generated by gpt-5.6-luna for the original question.
is_underspecified
Boolean: true if at least one source annotation has type multipleQAs; false otherwise (singleAnswer).
nq_answer
Original list of Natural Questions answers… See the full description on the dataset page: https://huggingface.co/datasets/zykov/AmbigQA-clarifications.
