CoolFace
Datasetpublic

undertheseanlp/sentence-segmentation-1

Sentence Segmentation Test set for evaluating and improving Vietnamese sentence boundary detection (sent_tokenize) in underthesea. Problem The current PunktSentenceTokenizer in underthesea fails on several Vietnamese-specific patterns, primarily in legal text where article titles are merged with sentence bodies without punctuation boundaries. Current Results Category Total Correct Accuracy title_content_merge 38 0 0.0% repeated_title 13… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/sentence-segmentation-1.

sourceHugging Faceupdated 8mo agoView on Hugging Face
0likes19downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
undertheseanlp/sentence-segmentation-1 · CoolFace