CoolFace
Datasetpublic

pralia-labs/canoncite

CANONCITE A benchmark for exact-ID citation attribution and abstention over ten public-domain canonical corpora, in four scripts across four religious traditions plus Tamil ethical literature and Indian constitutional law. Every question is posed three ways: in English, in Hindi, and in the corpus's own native script. What this measures, and why it is not the usual thing Most attribution benchmarks ask whether a generated claim is supported by a retrieved passage.… See the full description on the dataset page: https://huggingface.co/datasets/pralia-labs/canoncite.

sourceHugging Facecc-by-4.0updated 25d agoView on Hugging Face
0likes42downloads
3 commits on main
e759f0a25d ago

Point the code link at pralia-labs/canoncite

pralia-labs
3c5734225d ago

CANONCITE v0: 10 corpora, 188,557 units, 622 items

pralia-labs
a854e5725d ago

initial commit

pralia-labs