Sinhala
Datasets
All datasets matching “Sinhala”sinhala-flansentence_alignment_dataset-Sinhala-Tamil-English
Dataset summary
This is a gold-standard benchmark dataset for sentence alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. The aligned documents annotated in the dataset NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English had been considered to annotate the aligned sentences.
News Source
url
Army
https://www.army.lk/
Hiru
http://www.hirunews.lk
ITN
https://www.newsfirst.lk
Newsfirst
https://www.itnnews.lk… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English.document_alignment_dataset-Sinhala-Tamil-English
Dataset summary
This is a gold-standard benchmark dataset for document alignment, between Sinhala-English-Tamil languages.
Data had been crawled from the following news websites.
News Source
url
Army
https://www.army.lk/
Hiru
http://www.hirunews.lk
ITN
https://www.newsfirst.lk
Newsfirst
https://www.itnnews.lk
The aligned documents have been manually annotated.
Dataset
The folder structure for each news source is as follows.
army
|--Sinhala… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English.sinhala_dataset_vsinhala-22gb-cleaned-datasetsinhala-tts-dataset
Sinhala TTS Dataset
Clean, segmented single-speaker Sinhala speech from the "Unlimited History" YouTube series by @sunchare. Built for TTS fine-tuning (F5-TTS, VITS, etc.).
Dataset Versions
cc_v1 — Full dataset (63 videos)
Metric
Value
Utterances
22,441
Train / Val
21,319 / 1,122
Hours
23.23h
Mean duration
3.73s
Duration range
3.0s – 19.74s
Sample rate
22,050 Hz
Videos processed
63
Avg keep rate
84.4%… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset.
