CoolFace
Datasetpublic

Minuri/sinhala-corpus-madlad400

Sinhala Raw Sentences - MADLAD-400 Raw Sinhala sentences extracted and sentence-split from the allenai/MADLAD-400 dataset. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset. Dataset Structure Column Description text Raw Sinhala sentence source Source identifier (madlad) Split Rows train 7,281,026 Pipeline Position allenai/MADLAD-400 → this repo → Minuri/madlad_cleaned_version… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-madlad400.

sourceHugging Facecc-by-sa-4.0updated 6mo agoView on Hugging Face
0likes37downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Minuri/sinhala-corpus-madlad400 · CoolFace