trungbb8/vietnamese-news-copus-segmented
Dataset Card for vietnamese-news-copus-segmented Dataset Summary This dataset is a refined collection of Vietnamese news articles, originally sourced from ademax/binhvq-news-corpus. It has been processed through a specialized pipeline for cleaning, normalization, and word segmentation. It is ideal for training Vietnamese Language Models (LLMs), word embeddings, or text classification tasks. Original Source: ademax/binhvq-news-corpus Language: Vietnamese (vi)… See the full description on the dataset page: https://huggingface.co/datasets/trungbb8/vietnamese-news-copus-segmented.
0207
