CoolFace
Datasetpublic

sobir-hf/tajik-text-segmentation

This dataset contains texts in Tajik language with sentence annotations. It can be used to train and evaluate sentence-wise text segmentation algorithms. The dataset contains more than 100 short and long texts and more than 3000 annotated sentences. The texts were carefully selected from different catergories such as news, articles, novels, classical texts, poetry, and religious texts. It deliberately contains more of "hard" passages where splitting them by period "." characters would result… See the full description on the dataset page: https://huggingface.co/datasets/sobir-hf/tajik-text-segmentation.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
1likes64downloads
15 commits on main
f17f7373y ago

strip sentence ends when parsing

sobir-hf
76591253y ago

corrected more annotation mistakes

sobir-hf
e35a36d3y ago

fixed some annotation mistakes

sobir-hf
f436b4e3y ago

bugfix where # was interfering in parsing

sobir-hf
518362d3y ago

updated format and data

sobir-hf
e00c03d3y ago

Update README.nd

sobir-hf
52406a83y ago

converted the files to jsonl

sobir-hf
2432e743y ago

Update tajik-text-segmentation.py

sobir-hf
edd2ab83y ago

Update tajik-text-segmentation.py

sobir-hf
edc8edd3y ago

Create README.md

sobir-hf
cda116a3y ago

Update tajik-text-segmentation.py

sobir-hf
7b4dd763y ago

Update tajik-text-segmentation.py

sobir-hf
74f14473y ago

Upload tajik-text-segmentation.py

sobir-hf
453bed43y ago

Uploaded data files and parser script

sobir-hf
5be27633y ago

initial commit

sobir-hf