undertheseanlp/sentence-segmentation-1
Sentence Segmentation Test set for evaluating and improving Vietnamese sentence boundary detection (sent_tokenize) in underthesea. Problem The current PunktSentenceTokenizer in underthesea fails on several Vietnamese-specific patterns, primarily in legal text where article titles are merged with sentence bodies without punctuation boundaries. Current Results Category Total Correct Accuracy title_content_merge 38 0 0.0% repeated_title 13… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/sentence-segmentation-1.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face