CoolFace
Datasetpublic

trungbb8/vietnamese-news-copus-segmented

Dataset Card for vietnamese-news-copus-segmented Dataset Summary This dataset is a refined collection of Vietnamese news articles, originally sourced from ademax/binhvq-news-corpus. It has been processed through a specialized pipeline for cleaning, normalization, and word segmentation. It is ideal for training Vietnamese Language Models (LLMs), word embeddings, or text classification tasks. Original Source: ademax/binhvq-news-corpus Language: Vietnamese (vi)… See the full description on the dataset page: https://huggingface.co/datasets/trungbb8/vietnamese-news-copus-segmented.

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes207downloads

trungbb8/vietnamese-news-copus-segmented · main · files are served by the source, never re-hosted here