gt-csse/false-citation-bench
False Citation Bench False Citation Bench is a compact evaluation and inspection dataset for false or misleading case citations in legal documents. It contains 26 source documents, their PDFs, and manually reviewed citation annotations grounded in the local text extraction. Dataset contents The repository has one matching document in each directory: documents_txt/{index}__{case-name}__{filing}.txt documents_pdf/{index}__{case-name}__{filing}.pdf… See the full description on the dataset page: https://huggingface.co/datasets/gt-csse/false-citation-bench.
Note the provenance under which these filings are redistributed
Require complete identifiers, rename validation set, add CourtListener comparison corpus
Rename validation set, rewrite derived cards, formalize minimum sufficient case identifier, add CourtListener comparison corpus
Rename identity set to validation-courtlistener-heuristics; rewrite derived data cards
Add docket records to extraction set; document derived sets
Add docket records to extraction set; document derived sets
Add docket records to extraction set; rename locator_span to span
Add docket records to extraction set; rename locator_span to span
Mask self-reference docket numbers; add derived extraction and identity sets
Upload folder using huggingface_hub
Remove legacy numeric filenames
Rename files for human-readable dataset navigation
Add dataset card
Upload false citation benchmark dataset
initial commit
