CoolFace
Modelpublic

dewdev/language_detection

sourceHugging Facemitupdated 2y agoView on Hugging Face
1likes47downloads
Model Card

This is a clone of https://huggingface.co/alexneakameni/language_detection with onnx format

Language Detection Model

A BERT-based language detection model trained on hac541309/open-lid-dataset, which includes 121 million sentences across 200 languages. This model is optimized for fast and accurate language identification in text classification tasks.

Model Details

  • Architecture: BertForSequenceClassification
  • Hidden Size: 384
  • Number of Layers: 4
  • Attention Heads: 6
  • Max Sequence Length: 512
  • Dropout: 0.1
  • Vocabulary Size: 50,257

Training Process

  • Dataset:
  • Used the open-lid-dataset
  • Split into train (90%) and test (10%)
  • Tokenizer: A custom BertTokenizerFast with special tokens for [UNK], [CLS], [SEP], [PAD], [MASK]
  • Hyperparameters:
  • Learning Rate: 2e-5
  • Batch Size: 256 (training) / 512 (testing)
  • Epochs: 1
  • Scheduler: Cosine
  • Trainer: Leveraged the Hugging Face Trainer API with Weights & Biases for logging

Evaluation

The model was evaluated on the test split. Below are the overall metrics:

  • Accuracy: 0.969466
  • Precision: 0.969586
  • Recall: 0.969466
  • F1 Score: 0.969417

Detailled evaluation (Size is the number of languages supported)

ScriptSupportPrecisionRecallF1 ScoreSize
Arab8192190.90380.90140.902321
Latn79247040.96780.96630.9670125
Ethi1444030.99670.99640.99662
Beng1639830.99490.99350.99423
Deva4238950.94950.93260.940510
Cyrl8319490.98990.98830.989112
Tibt356830.99250.99300.99272
Grek1311550.99840.99900.99871
Gujr869120.999990.99990.999951
Hebr1005300.99660.99950.99812
Armn672030.99990.99980.99981
Jpan880040.99830.99870.99851
Knda671700.99990.99980.99991
Geor707690.999970.99980.99991
Khmr397081.00000.99970.99991
Hang1085090.99970.99990.99981
Laoo293890.99990.99990.99991
Mlym684180.999960.99990.99991
Mymr1008570.99990.99920.99952
Orya449760.99950.99980.99961
Guru671060.999990.99990.99991
Olck222791.00000.99910.99951
Sinh674921.00000.99980.99991
Taml763730.999970.99990.99991
Tfng413250.85120.82460.82472
Telu623870.999970.99990.99991
Thai838200.999950.99980.99991
Hant1527230.99450.99540.99492
Hans926890.98930.98700.98821

A detailed per-script classification report is also provided in the repository for further analysis.


How to Use

You can quickly load and run inference with this model using the Transformers pipeline:

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification, pipeline

tokenizer = AutoTokenizer.from_pretrained("alexneakameni/language_detection")
model = AutoModelForSequenceClassification.from_pretrained("alexneakameni/language_detection")

language_detection = pipeline("text-classification", model=model, tokenizer=tokenizer)

text = "Hello world!"
predictions = language_detection(text)
print(predictions)

This will output the predicted language code or label with the corresponding confidence score.


Note: The model’s performance may vary depending on text length, language variety, and domain-specific vocabulary. Always validate results against your own datasets for critical applications.

For more information, see the repository documentation.

Thank you for using this model—feedback and contributions are welcome!