CoolFace
Datasetpublic

DS4H-ICTU/yat-ner-dataset

Yambeta Named Entity Recognition (NER) Dataset Dataset Description This dataset was developed for Named Entity Recognition (NER) tasks in Yambeta, a Bantu language from Cameroon. The dataset contains sentences from the Yambeta Bible text corpus, annotated with named entities such as persons, locations, and organizations. Developed by: DS4H-ICTU Research Group in cooperation with the Yambeta Bible Project. Language(s): Yambeta (Bantu language from Cameroon)… See the full description on the dataset page: https://huggingface.co/datasets/DS4H-ICTU/yat-ner-dataset.

sourceHugging Faceupdated 8mo agoView on Hugging Face
0likes12downloads
Dataset Card

Yambeta Named Entity Recognition (NER) Dataset

Dataset Description

This dataset was developed for Named Entity Recognition (NER) tasks in Yambeta, a Bantu language from Cameroon. The dataset contains sentences from the Yambeta Bible text corpus, annotated with named entities such as persons, locations, and organizations.

  • —Developed by: DS4H-ICTU Research Group in cooperation with the Yambeta Bible Project.
  • —Language(s): Yambeta (Bantu language from Cameroon)
  • —License: Apache 2.0 (or specify if different)
  • —Dataset Type: Annotated Text (NER)

Dataset Sources

  • —Source Text: Yambeta Bible text corpus (final_dataset.xlsx)
  • —Number of Samples: {total_samples} samples
  • —Named Entities: {totalentities} unique entities across {entitytypes} categories (persons, locations, organizations)

Uses

  • —Direct Use: This dataset can be used for training, fine-tuning, and evaluating NER models in Yambeta.
  • —Downstream Use: Models trained on this dataset can be utilized for applications like translation, entity recognition, or information extraction.

Bias, Risks, and Limitations

  • —Biases: The dataset is extracted from a religious text and may not represent a comprehensive set of Yambeta language entities.
  • —Out-of-Scope Use: The dataset may not generalize well to non-religious texts in Yambeta.

Dataset Structure

  • —Format: JSON or DatasetDict with tokens and ner_tags fields.
  • —Train/Validation/Test Split:
  • —Train set: 80%
  • —Validation set: 10%
  • —Test set: 10%

Citation

If you use this dataset in your work, please cite it using the following format:

@misc{yambeta_ner_dataset,
  title = {Yambeta NER Dataset},
  author = {Dr.-Ing. Philippe Tamla},
  year = {2024},
  publisher = {Hugging Face},
  url = {https://huggingface.co/DS4H-ICTU/yat-ner-datasets}
}

Contact Information

For more information, contact the developers at: philiptamla@gmail.com