CoolFace
Datasetpublic

sello-ralethe/Knowledge_Base_Projection

Dataset Summary This is a cross-lingual knowledge base and question answering dataset for four low-resource South African languages: isiZulu, isiXhosa, Sepedi, and SeSotho. The dataset includes: Parallel text corpora for alignment Projected knowledge bases from ConceptNet and DBpedia Verbalized Triples Translated question-answer pairs The dataset was created using LeNS-Align, a novel cross-lingual mapping technique that combines lexical alignment, named entity recognition, and semantic… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/Knowledge_Base_Projection.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes71downloads
README.md23 linesDownload Raw Back to root
1---2license: apache-2.03---4Dataset Summary5 6 This is a cross-lingual knowledge base and question answering dataset for four low-resource South African languages: isiZulu, isiXhosa, Sepedi, and SeSotho. The dataset includes:7 81. Parallel text corpora for alignment92. Projected knowledge bases from ConceptNet and DBpedia103. Verbalized Triples114. Translated question-answer pairs12 13The dataset was created using LeNS-Align, a novel cross-lingual mapping technique that combines lexical alignment, named entity recognition, and semantic alignment to project knowledge from English to low-resource languages.14 15Languages16 17- isiZulu (zul)18- isiXhosa (xho)19- Sepedi (nso)20- SeSotho (sot)21- English (eng)22 23