CoolFace
Datasetpublic

Minuri/sinhala-corpus-madlad400

Sinhala Raw Sentences - MADLAD-400 Raw Sinhala sentences extracted and sentence-split from the allenai/MADLAD-400 dataset. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset. Dataset Structure Column Description text Raw Sinhala sentence source Source identifier (madlad) Split Rows train 7,281,026 Pipeline Position allenai/MADLAD-400 → this repo → Minuri/madlad_cleaned_version… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-madlad400.

sourceHugging Facecc-by-sa-4.0updated 6mo agoView on Hugging Face
0likes37downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

Minuri/sinhala-corpus-madlad400 · main · files are served by the source, never re-hosted here