rafmacalaba/datause-extracted-sample
Probe-review sample of rafmacalaba/datause-extracted-sample Stratified review slice: whole rows sampled per origin until each specificity reaches ~150 spans (seed 0; small origins contribute all they have), plus entity-less calibration rows round-robined across neg_class shapes. Rows keep their original split values. Score with training/score_extract_probe.py --splits sample. origin rows negatives named / descriptive / vague spans fcv_pads_east_africa 389 60 171 / 187… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-extracted-sample.
Probe-review sample of rafmacalaba/datause-extracted-sample
Stratified review slice: whole rows sampled per origin until each specificity reaches ~150 spans (seed 0; small origins contribute all they have), plus entity-less calibration rows round-robined across neg_class shapes. Rows keep their original split values. Score with training/scoreextractprobe.py --splits sample.
Probe scores
Every span carries probe_score (sigmoid head output ~= P(genuine data-use mention)) from rafmacalaba/gliner_datause_probe (frozen rafmacalaba/gliner_datause encoder, context_radius=64, threshold=None). Extractor proposes, head disposes — see probe.json.
Review split (gliner2_review.jsonl, 2026-09-06)
421 passages / 473 spans from the annotator boundary set (190) and JDC operational review (283), added as origins annotator190 / jdc283 with split: review. Labels are the reconciled v2.4 gold (3 fresh single-span runs + Sonnet tiebreak + human adjudication; see rafmacalaba/datause-v24-sample-verdicts, doctrinefindings.md). Span entries carry `key`, `reviewtier`, and char offsets.
