CoolFace
Modelpublic

no-name-research/multilingual-bert-spatial-relations-classifier

sourceHugging Facecc-by-nc-4.0updated 1y agoView on Hugging Face
0likes4downloads
Model Card

bert-base-multilingual-cased-classification-relation

<!-- Provide a quick summary of what the model is/does. -->

This model is designed to classify spatial relations recognized from geographic encyclopedia articles. It is a fine-tuned version of the bert-base-multilingual-cased model. It has been trained on no-name-research/no-name-dataset, a manually annotated subset of the French Encyclopédie ou dictionnaire raisonné des sciences des arts et des métiers par une société de gens de lettres (1751-1772) edited by Diderot and d'Alembert (provided by the ARTFL Encyclopédie Project).

Model Description

<!-- Provide a longer summary of what this model is. -->

  • Authors: xxxxxxxxxx
  • Model type: Text classification
  • Repository: xxxxxxxxxxx
  • Language(s) (NLP): French
  • License: cc-by-nc-4.0

Class labels

The tagset is as follows:

  • Adjacency:
  • Crosses:
  • Distance-Orientation:
  • Inclusion:
  • Movement:
  • Other:

Dataset

The model was trained using the no-name-research/no-name-dataset dataset. The dataset is splitted into train, validation and test sets which have the following distribution of entries among classes:

TrainValidationTest
Adjacency4985975
Crosses3975029
Distance-Orientation1,065163115
Inclusion1,319131156
Movement1841535
Other1953042

Evaluation

  • Overall weighted-average model performances
PrecisionRecallF-score
0.920.920.92
  • Model performances (Test set)
PrecisionRecallF-scoreSupport
Adjacency0.850.840.8575
Crosses0.780.860.8229
Distance-Orientation0.930.990.96115
Inclusion0.970.980.97156
Movement0.890.690.7735
Other0.950.880.9142

How to Get Started with the Model

Use the code below to get started with the model.

python
import torch
from transformers import pipeline, AutoTokenizer, AutoModelForSequenceClassification
device = torch.device("mps" if torch.backends.mps.is_available() else ("cuda" if torch.cuda.is_available() else "cpu"))

ner = pipeline("token-classification", model="no-name-research/camembert-token-classification", aggregation_strategy="simple", device=device)
relation_classifier = pipeline("text-classification", model="no-name-research/multilingual-bert-spatial-relations-classifier", truncation=True, device=device)

def get_context(text, span, ngram_context_size=5):
    word = span["word"]
    start = span["start"]
    end = span["end"]
    label = span["entity_group"]

    # Extract context
    previous_text = text[:start].strip()
    next_text = text[end:].strip()
    previous_words = previous_text.split()[-ngram_context_size:]
    next_words = next_text.split()[:ngram_context_size]

    # Build context string
    context = f"[{word}]: {' '.join(previous_words)} {word} {' '.join(next_words)}"
    return word, context, label

content = "WINCHESTER, (Géog. mod.) ou plutôt Wintchester, ville d'Angleterre, capitale du Hampshire, sur le bord de l'Itching, à dix-huit milles au sud-est de Salisbury, & à soixante sud-ouest de Londres. Long. 16. 20. latit. 51. 3."

spans = ner(content)
for span in spans:
    if span['entity_group'] == 'Relation':
        word, context, label = get_context(content, span, ngram_context_size=5)
        print(f"Relation: {word}")

        label = relation_classifier(context)
        print(f"Predicted label: {label}")


# Output
Relation: sur le bord de
Predicted label: [{'label': 'Crosses', 'score': 0.9778845906257629}]
Relation: à dix-huit milles au sud-est de
Predicted label: [{'label': 'Distance-Orientation', 'score': 0.9959626793861389}]
Relation: à soixante sud-ouest de
Predicted label: [{'label': 'Distance-Orientation', 'score': 0.9963018894195557}]

Bias, Risks, and Limitations

<!-- This section is meant to convey both technical and sociotechnical limitations. -->

This model was trained entirely on French encyclopaedic entries classified as Geography and will likely not perform well on text in other languages or other corpora.