CoolFace
Datasetpublic

sello-ralethe/Knowledge_Base_Projection

Dataset Summary This is a cross-lingual knowledge base and question answering dataset for four low-resource South African languages: isiZulu, isiXhosa, Sepedi, and SeSotho. The dataset includes: Parallel text corpora for alignment Projected knowledge bases from ConceptNet and DBpedia Verbalized Triples Translated question-answer pairs The dataset was created using LeNS-Align, a novel cross-lingual mapping technique that combines lexical alignment, named entity recognition, and semantic… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/Knowledge_Base_Projection.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes71downloads
Dataset Card

Dataset Summary

This is a cross-lingual knowledge base and question answering dataset for four low-resource South African languages: isiZulu, isiXhosa, Sepedi, and SeSotho. The dataset includes:

  1. 1.Parallel text corpora for alignment
  2. 2.Projected knowledge bases from ConceptNet and DBpedia
  3. 3.Verbalized Triples
  4. 4.Translated question-answer pairs

The dataset was created using LeNS-Align, a novel cross-lingual mapping technique that combines lexical alignment, named entity recognition, and semantic alignment to project knowledge from English to low-resource languages.

Languages

  • isiZulu (zul)
  • isiXhosa (xho)
  • Sepedi (nso)
  • SeSotho (sot)
  • English (eng)