sentence-alignment
sentence_alignment_dataset-Sinhala-Tamil-English
Dataset summary
This is a gold-standard benchmark dataset for sentence alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. The aligned documents annotated in the dataset NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English had been considered to annotate the aligned sentences.
News Source
url
Army
https://www.army.lk/
Hiru
http://www.hirunews.lk
ITN
https://www.newsfirst.lk
Newsfirst
https://www.itnnews.lk… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English.sentence-alignment-tib-eng
Dataset Card for "sentence-alignment-tib-eng"
More Information needed
sentence-alignment-merged-postcorrection
Dataset Card for "sentence-alignment-merged-postcorrection"
More Information needed
SentenceAlignment
SentenceAlignment
tags: correspondence, machine learning, contextual
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'SentenceAlignment' dataset comprises paragraphs and their associated claims that are expressed as exact sentences found within the paragraph. This dataset can be used for training machine learning models to identify and extract claims directly from text passages. It is relevant for natural language processing… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/SentenceAlignment.
