CoolFace
Datasetpublic

amritha27/cl3410-phase1

CL3410 Phase 1 — Malayalam and Assamese language-model corpora Two independently built pretraining corpora with their own tokenizers: Malayalam as the higher-resource language and Assamese as the lower-resource one. Nothing is shared between them — separate sources, separate cleaning thresholds, separate vocabularies, separate models. Only the language-agnostic pipeline code is common, parameterised per language. Everything here was collected and cleaned for this project. No… See the full description on the dataset page: https://huggingface.co/datasets/amritha27/cl3410-phase1.

sourceHugging Faceotherupdated 10d agoView on Hugging Face
0likes160downloads
50 commits on main
e3d4d9a10d ago

Upload folder using huggingface_hub

amritha27
4e4c8ba10d ago

Upload folder using huggingface_hub

amritha27
26088ea10d ago

Upload folder using huggingface_hub

amritha27
1776cb810d ago

Upload folder using huggingface_hub

amritha27
52828fc14d ago

Upload scorers/ml_quality.npz with huggingface_hub

amritha27
88d29e814d ago

Upload scorers/ml_ppl.npz with huggingface_hub

amritha27
ec1ca9f14d ago

Upload scorers/as_quality.npz with huggingface_hub

amritha27
f1f09d714d ago

Upload scorers/as_ppl.npz with huggingface_hub

amritha27
a8f68a721d ago

Upload checkpoints/as/step_0007350.pt with huggingface_hub

amritha27
d96177221d ago

Upload checkpoints/as/step_0007250.pt with huggingface_hub

amritha27
96d713721d ago

Upload checkpoints/as/step_0007000.pt with huggingface_hub

amritha27
55d6d6321d ago

Upload checkpoints/as/final.pt with huggingface_hub

amritha27
4c0694221d ago

Upload checkpoints/ml/step_0007631.pt with huggingface_hub

amritha27
a013e3821d ago

Upload checkpoints/ml/step_0007500.pt with huggingface_hub

amritha27
80bc0b521d ago

Upload checkpoints/ml/step_0007250.pt with huggingface_hub

amritha27
cf8a57621d ago

Upload checkpoints/ml/final.pt with huggingface_hub

amritha27
2d6f2d221d ago

Upload checkpoints/as/verify.json with huggingface_hub

amritha27
509979721d ago

Upload checkpoints/as/metrics.csv with huggingface_hub

amritha27
f21195321d ago

Upload checkpoints/as/run_config.json with huggingface_hub

amritha27
0dbefdb21d ago

Upload checkpoints/as/best.pt with huggingface_hub

amritha27
259010921d ago

Upload checkpoints/ml/verify.json with huggingface_hub

amritha27
4b6ae9721d ago

Upload checkpoints/ml/metrics.csv with huggingface_hub

amritha27
f08b4ff21d ago

Upload checkpoints/ml/run_config.json with huggingface_hub

amritha27
326f86b21d ago

Upload checkpoints/ml/best.pt with huggingface_hub

amritha27
abd619e1mo ago

Add token rank-frequency figure

amritha27
397b7b61mo ago

Add language-selection justification

amritha27
6a28f141mo ago

Expand filter documentation

amritha27
b9ae7141mo ago

Fix coherence figures and document the inert filter

amritha27
ad346f21mo ago

Add vocabulary-choice justification and figure

amritha27
f33cd3a1mo ago

Clarify held-out split usage

amritha27
e679b611mo ago

Measure the token target against the train split

amritha27
7a722ba1mo ago

Add ocr.tar

amritha27
b42cde61mo ago

Add manual.tar

amritha27
8bfa2e71mo ago

Add harvested_docs.jsonl

amritha27
084492f1mo ago

Sync 2 files (16)

amritha27
6c0bddc1mo ago

Sync 1 files (15)

amritha27
6231b271mo ago

Sync 1 files (14)

amritha27
e82a9e91mo ago

Sync 1 files (13)

amritha27
5658b9e1mo ago

Sync 7 files (12)

amritha27
3d7f1271mo ago

Sync 1 files (11)

amritha27
44807961mo ago

Sync 1 files (10)

amritha27
525853b1mo ago

Sync 1 files (9)

amritha27
de7c0e81mo ago

Sync 1 files (8)

amritha27
7d7459a1mo ago

Sync 1 files (7)

amritha27
b0a47331mo ago

Sync 1 files (6)

amritha27
f9fde351mo ago

Sync 1 files (5)

amritha27
a53fa6d1mo ago

Sync 1 files (4)

amritha27
a216f1c1mo ago

Sync 1 files (3)

amritha27
2b978e21mo ago

Sync 1 files (2)

amritha27
34488d11mo ago

Sync 1 files (1)

amritha27