pralia-labs/canoncite
CANONCITE A benchmark for exact-ID citation attribution and abstention over ten public-domain canonical corpora, in four scripts across four religious traditions plus Tamil ethical literature and Indian constitutional law. Every question is posed three ways: in English, in Hindi, and in the corpus's own native script. What this measures, and why it is not the usual thing Most attribution benchmarks ask whether a generated claim is supported by a retrieved passage.… See the full description on the dataset page: https://huggingface.co/datasets/pralia-labs/canoncite.
042
Point the code link at pralia-labs/canoncite
CANONCITE v0: 10 corpora, 188,557 units, 622 items
initial commit
