CoolFace
Datasetpublic

ananddey/assamese-wiki-corpus

Assamese Wiki Corpus (ananddey/assamese-wiki-corpus) A clean, large scale Assamese text corpus spanning wiki articles, literary works, dictionary entries, and quotations, curated for language model pre training, fine tuning, and NLP research. Total characters: 65,899,040Approximate tokens : 16,474,760 (16.5M) Fields Field Type Description id int64 Wikimedia page ID title string Page title text string Cleaned plain text content source class_label… See the full description on the dataset page: https://huggingface.co/datasets/ananddey/assamese-wiki-corpus.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
1likes25downloads
settings

This repository belongs to ananddey on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameassamese-wiki-corpus
visibilitypublic
licencecc-by-sa-4.0
gatedno
ownerananddey
Account settings
ananddey/assamese-wiki-corpus · CoolFace