CoolFace
Datasetpublic

Minuri/sinhala-corpus-madlad400

Sinhala Raw Sentences - MADLAD-400 Raw Sinhala sentences extracted and sentence-split from the allenai/MADLAD-400 dataset. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset. Dataset Structure Column Description text Raw Sinhala sentence source Source identifier (madlad) Split Rows train 7,281,026 Pipeline Position allenai/MADLAD-400 → this repo → Minuri/madlad_cleaned_version… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-madlad400.

sourceHugging Facecc-by-sa-4.0updated 6mo agoView on Hugging Face
0likes37downloads
settings

This repository belongs to Minuri on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namesinhala-corpus-madlad400
visibilitypublic
licencecc-by-sa-4.0
gatedno
ownerMinuri
Account settings
Minuri/sinhala-corpus-madlad400 · CoolFace