CoolFace
Modelpublic

alexgoldberg/hebrew-manuscript-joint-ner-v2

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes28downloads
Model Card

Hebrew Manuscript Joint NER v2

This repository contains the MHM Pipeline person NER model. The current checkpoint is the role-aware v3 replacement for the earlier custom two-head checkpoint, while keeping the same repository and bundle name for compatibility.

The model is a DictaBERT token-classification checkpoint that predicts BIO labels with the person role encoded directly in the tag:

  • —AUTHOR
  • —TRANSCRIBER
  • —OWNER
  • —CENSOR
  • —TRANSLATOR
  • —COMMENTATOR

Evaluation

Held-out v3 test split, 904 items:

MetricScore
strict span + role F10.8031
strict precision0.7888
strict recall0.8180
name-only F10.8665
role accuracy when name matched0.9269

Per-role strict span+role F1:

RoleF1
AUTHOR0.8678
CENSOR0.8830
COMMENTATOR0.5185
OWNER0.7330
TRANSCRIBER0.8112
TRANSLATOR0.9072

Usage

python
from transformers import AutoModelForTokenClassification, AutoTokenizer

repo_id = "alexgoldberg/hebrew-manuscript-joint-ner-v2"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForTokenClassification.from_pretrained(repo_id)

In MHM Pipeline, use ner.inference_pipeline.JointNERPipeline. It preserves the legacy output schema:

python
from ner.inference_pipeline import JointNERPipeline

pipeline = JointNERPipeline("alexgoldberg/hebrew-manuscript-joint-ner-v2")
entities = pipeline.process_text("הספר נכתב על ידי משה בן יעקב.")

Example output:

json
[
  {
    "person": "משה בן יעקב",
    "role": "TRANSCRIBER",
    "confidence": 0.9918,
    "model_confidence": 0.9918,
    "start": 17,
    "end": 28
  }
]

Notes

The previous custom checkpoint can be recovered from the Hub commit history. This version intentionally replaces keyword-based role classification with neural role-aware BIO labels.