sello-ralethe/Knowledge_Base_Projection
Dataset Summary This is a cross-lingual knowledge base and question answering dataset for four low-resource South African languages: isiZulu, isiXhosa, Sepedi, and SeSotho. The dataset includes: Parallel text corpora for alignment Projected knowledge bases from ConceptNet and DBpedia Verbalized Triples Translated question-answer pairs The dataset was created using LeNS-Align, a novel cross-lingual mapping technique that combines lexical alignment, named entity recognition, and semantic… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/Knowledge_Base_Projection.
Dataset Summary
This is a cross-lingual knowledge base and question answering dataset for four low-resource South African languages: isiZulu, isiXhosa, Sepedi, and SeSotho. The dataset includes:
- Parallel text corpora for alignment
- Projected knowledge bases from ConceptNet and DBpedia
- Verbalized Triples
- Translated question-answer pairs
The dataset was created using LeNS-Align, a novel cross-lingual mapping technique that combines lexical alignment, named entity recognition, and semantic alignment to project knowledge from English to low-resource languages.
Languages
- isiZulu (zul)
- isiXhosa (xho)
- Sepedi (nso)
- SeSotho (sot)
- English (eng)
