amritha27/cl3410-phase1
CL3410 Phase 1 — Malayalam and Assamese language-model corpora Two independently built pretraining corpora with their own tokenizers: Malayalam as the higher-resource language and Assamese as the lower-resource one. Nothing is shared between them — separate sources, separate cleaning thresholds, separate vocabularies, separate models. Only the language-agnostic pipeline code is common, parameterised per language. Everything here was collected and cleaned for this project. No… See the full description on the dataset page: https://huggingface.co/datasets/amritha27/cl3410-phase1.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face