ganga4364/tibetan-metadata-detector
Tibetan Metadata Detector Dataset RoBERTa sliding-window training data for Tibetan title and author span detection from BDRC outliner exports. Contents (windows config) Balanced window splits (fixed window-relative BIO labeling, O-only subsampling, author oversampling): Split Description train ~89% of documents (stratified) validation ~1% (small val for fast eval) test ~10% Balancing applied before split: O-only windows capped at 2×… See the full description on the dataset page: https://huggingface.co/datasets/ganga4364/tibetan-metadata-detector.
0104
