Ahren09/SARA-QASPER
SARA QASPER (reformatted) Reformatted QASPER data used by SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression (ACL 2026, arXiv:2507.05633). Code: Ahren09/SARA. The SARA Quick Start (python -m src.data.make_qasper_splits) downloads this dataset automatically; you can also load it directly: from datasets import load_dataset qa = load_dataset("Ahren09/SARA-QASPER", "qa") # train / test align =… See the full description on the dataset page: https://huggingface.co/datasets/Ahren09/SARA-QASPER.
SARA QASPER (reformatted)
Reformatted QASPER data used by SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression (ACL 2026, arXiv:2507.05633). Code: Ahren09/SARA.
The SARA Quick Start (python -m src.data.make_qasper_splits) downloads this dataset automatically; you can also load it directly:
from datasets import load_dataset
qa = load_dataset("Ahren09/SARA-QASPER", "qa") # train / test
align = load_dataset("Ahren09/SARA-QASPER", "compression_alignment") # train / validationConfigs
qa (default)
One record per answerable QASPER question. train (2,321 rows over the official train papers) and test (1,312 rows, official test papers).
Models trained on this data answer with chain-of-thought followed by the final short answer wrapped in tags: ... reasoning ... <answer>short answer</answer>; evaluation extracts the tagged span and scores it against answer.
compression_alignment
Text snippets from QASPER paper bodies used for the SARA projector-alignment warm-up (the projector learns to reconstruct a document from its semantic compression vector before QA fine-tuning). train (22,111 rows) and validation (200 rows); each record is {"text": "<document text snippet>"}.
Provenance
- Derived from `allenai/qasper` (Dasigi et al., NAACL 2021), released under CC BY 4.0. This derivative is released under the same license with attribution.
contextwas built from each paper's title/abstract/sections and BM25-ranked per question;BIBREFcitation markers were stripped.- Earlier revisions of this dataset carried an additional LLM-rewritten
answer_reformattedfield; it has been removed — the short goldansweris the single reference. The old revision remains available in this repo's git history.
Checksums (sha256)
0ec3d1bbab2f85341a9432b263ca90669f5cfd45bb2163170cedddf8ff60742c QASPER_train.jsonl
a26afe8a11e350e5a860c83b75ee4938b2b89f19dc4883cea063dd90642bc1e1 QASPER_test.jsonl
1d6386a84408127a20f01db92f8b9a73ec9499d884d0cab3c2c9f07c4494aa12 QASPER_compression_alignment_train.jsonl
9de81c3cbf42aa4b4001a57762c37f981eb2ae0c5b468c3bfebffe65438f8a0f QASPER_compression_alignment_dev.jsonlCitation
@inproceedings{jin2025sara,
title={SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression},
author={Jin, Yiqiao and Sharma, Kartik and Rakesh, Vineeth and Dou, Yingtong and Pan, Menghai and Das, Mahashweta and Kumar, Srijan},
booktitle={ACL},
year={2026}
}
@inproceedings{dasigi2021dataset,
title={A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers},
author={Dasigi, Pradeep and Lo, Kyle and Beltagy, Iz and Cohan, Arman and Smith, Noah A. and Gardner, Matt},
booktitle={NAACL},
year={2021}
}