amritha27/cl3410-phase1
CL3410 Phase 1 — Malayalam and Assamese language-model corpora Two independently built pretraining corpora with their own tokenizers: Malayalam as the higher-resource language and Assamese as the lower-resource one. Nothing is shared between them — separate sources, separate cleaning thresholds, separate vocabularies, separate models. Only the language-agnostic pipeline code is common, parameterised per language. Everything here was collected and cleaned for this project. No… See the full description on the dataset page: https://huggingface.co/datasets/amritha27/cl3410-phase1.
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Upload scorers/ml_quality.npz with huggingface_hub
Upload scorers/ml_ppl.npz with huggingface_hub
Upload scorers/as_quality.npz with huggingface_hub
Upload scorers/as_ppl.npz with huggingface_hub
Upload checkpoints/as/step_0007350.pt with huggingface_hub
Upload checkpoints/as/step_0007250.pt with huggingface_hub
Upload checkpoints/as/step_0007000.pt with huggingface_hub
Upload checkpoints/as/final.pt with huggingface_hub
Upload checkpoints/ml/step_0007631.pt with huggingface_hub
Upload checkpoints/ml/step_0007500.pt with huggingface_hub
Upload checkpoints/ml/step_0007250.pt with huggingface_hub
Upload checkpoints/ml/final.pt with huggingface_hub
Upload checkpoints/as/verify.json with huggingface_hub
Upload checkpoints/as/metrics.csv with huggingface_hub
Upload checkpoints/as/run_config.json with huggingface_hub
Upload checkpoints/as/best.pt with huggingface_hub
Upload checkpoints/ml/verify.json with huggingface_hub
Upload checkpoints/ml/metrics.csv with huggingface_hub
Upload checkpoints/ml/run_config.json with huggingface_hub
Upload checkpoints/ml/best.pt with huggingface_hub
Add token rank-frequency figure
Add language-selection justification
Expand filter documentation
Fix coherence figures and document the inert filter
Add vocabulary-choice justification and figure
Clarify held-out split usage
Measure the token target against the train split
Add ocr.tar
Add manual.tar
Add harvested_docs.jsonl
Sync 2 files (16)
Sync 1 files (15)
Sync 1 files (14)
Sync 1 files (13)
Sync 7 files (12)
Sync 1 files (11)
Sync 1 files (10)
Sync 1 files (9)
Sync 1 files (8)
Sync 1 files (7)
Sync 1 files (6)
Sync 1 files (5)
Sync 1 files (4)
Sync 1 files (3)
Sync 1 files (2)
Sync 1 files (1)
