CoolFace
Datasetpublic

NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English

Dataset summary This is a gold-standard benchmark dataset for sentence alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. The aligned documents annotated in the dataset NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English had been considered to annotate the aligned sentences. News Source url Army https://www.army.lk/ Hiru http://www.hirunews.lk ITN https://www.newsfirst.lk Newsfirst… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English.

sourceHugging Faceupdated 3y agoView on Hugging Face
3likes656downloads

NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English · main · files are served by the source, never re-hosted here