CoolFace
Datasetpublic

BlossomsAI/vietnamese-corpus

Vietnamese Combined Corpus Dataset Statistics Total documents: {<15M:,} Wikipedia articles: {>1.3M:,} News articles: {>13M:,} Text documents: {>200K:,} Processing Details Processed using Apache Spark Minimum document length: {10} characters Text cleaning applied: HTML/special character removal Whitespace normalization URL removal Empty document filtering Data Format Each document has: 'text': The document content 'source':… See the full description on the dataset page: https://huggingface.co/datasets/BlossomsAI/vietnamese-corpus.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
8likes648downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
BlossomsAI/vietnamese-corpus · CoolFace