CoolFace
Modelpublic

BDRC/TiBLA-RTDETR

sourceHugging Faceagpl-3.0updated 15d agoView on Hugging Face
1likes81downloads
Model Card

TiBLA-RTDETR

Primary checkpoint of TiBLA (Tibetan Book Layout Analysis) — an RT-DETR-l detector for the page layout of modern Tibetan books (headers, text area, footers, footnotes).

  • Base model / provenance: RT-DETR-l (Ultralytics), fine-tuned on the leak-free v4 tam2col split of TiBLAD.
  • License: AGPL-3.0 (inherited from the Ultralytics RT-DETR weights).
  • Dataset: BDRC/TiBLAD
  • Paper: buda-base/papers (papers/2026-tibetan-book-layout) — arXiv link forthcoming
  • Code: github.com/buda-base/tibla
This checkpoint is seed 0. Across five training seeds the paper reports mean canonical F1 0.961 ± 0.009 (unified scorer, per-seed operating point); at the validation-selected operating point used in the table below this seed scores 0.952.

Task

A 4-class detector — header, text-area, footer, footnote — kept as four classes at training time. Evaluation folds them into a 3-class canonical scheme: header+footer are combined into one header-footer class (matched individually, merged losslessly afterwards), text-area is merged to a single page/column envelope as a post-processing step (two boxes only on genuine two-column pages), and footnote is left as-is. All numbers below are in that canonical space, on the leak-free TiBLAD v4 833-page test set, unified scorer (pycocotools bbox mAP 0.50:0.05:0.95; F1 by greedy IoU≥0.5 at the best-mean-F1 operating point).

Inference

python
# pip install ultralytics
from ultralytics import RTDETR

model = RTDETR("tibetan_book_layout.pt")
# recommended per-class confidence thresholds (see below); predict at the floor
res = model.predict("page.jpg", imgsz=1024, conf=0.25)[0]
TH = {0: 0.60, 1: 0.55, 2: 0.25, 3: 0.60}  # header / text-area / footnote / footer
for cls, conf, xywhn in zip(res.boxes.cls.tolist(), res.boxes.conf.tolist(),
                            res.boxes.xywhn.tolist()):
    if conf >= TH[int(cls)]:
        print(res.names[int(cls)], round(conf, 3), [round(v, 4) for v in xywhn])

A ready-made infer.py (batch, YOLO-format output) is included in this repo.

Recommended confidence thresholds (per-class max-F1 operating points): header/footer0.60, text-area0.55, footnote0.25. Footnote is deliberately kept low (recall-safe): the v4 test has only 38 footnote GT boxes, so a low threshold keeps recall near 1.0. Raising header/footer from 0.25 to 0.60 lifts precision +0.028 for a −0.014 recall cost; raising text-area from 0.25 to 0.55 (native) lifts precision +0.007 at no recall cost. If you prefer one global knob, the single validation-selected best-mean-F1 confidence is 0.64 (costs ≈0.008 mean F1 vs per-class tuning).

Evaluation (TiBLAD v4, 833-page test)

metric**TiBLA-RTDETR**TiBLA-PP-DocLayout-LTiBLA-RFDETR
licenseAGPL-3.0Apache-2.0Apache-2.0
base modelRT-DETR-l (Ultralytics)PP-DocLayout-L (PaddleOCR, RT-DETR-L)RF-DETR-L (Roboflow)
mean F1 (canonical 3-class)0.9520.9550.921
  header-footer F10.9540.9530.947
  text-area F10.9990.9980.996
  footnote F10.9020.9140.821
mean AP@0.500.9740.9590.925
mean AP@[0.50:0.95]0.7860.7810.667
shared-class mAP@[.50:.95] (DocLayNet-aligned)0.6500.6410.604
Hidden Trespass — header/footer0.0090.0040.021
Hidden Trespass — footnote0.0430.0370.178
COTe (Trespass)0.975 (0.001)0.978 (0.000)0.974 (0.002)
operating confidence0.640.610.47

"operating confidence" is the single global best-mean-F1 confidence, selected on the leak-free validation split and frozen for test (no test-set tuning). COCO AP rows are threshold-free (all detections above the fixed 0.05 floor).

Hidden Trespass = peripheral (header/footer/footnote) ground-truth area that survives in the actual OCR body crop C = E \ P, where E is the predicted text-area envelope and P is the union of the predicted peripheral boxes the pipeline subtracts; area-based, micro-averaged over the test set. Lower is better (less peripheral text bled into the OCR region). Formal definition in the paper.

Which checkpoint to pick

checkpointlicensemean F1shared mAPfootnote HT
TiBLA-RTDETR (primary)AGPL-3.00.9520.6500.043
TiBLA-PP-DocLayout-LApache-2.00.9550.6410.037
TiBLA-RFDETRApache-2.00.9210.6040.178

RT-DETR-l leads on mAP, shared-class mAP and the 5-seed mean F1 (0.961 ± 0.009), but its weights are AGPL-3.0 (Ultralytics). If you need a permissive license, PP-DocLayout-L is an Apache-2.0 match (statistically on par on F1); RF-DETR is a lighter PyTorch-native Apache-2.0 option.

Citation

bibtex
@misc{tibla2026,
  title        = {TiBLA: Tibetan Book Layout Analysis},
  author       = {Buddhist Digital Resource Center (BDRC)},
  year         = {2026},
  howpublished = {\url{https://github.com/buda-base/tibla}},
  note         = {arXiv link forthcoming}
}