Minuri/sinhala-corpus-culturax
Sinhala Raw Sentences - CulturaX Raw Sinhala sentences extracted and sentence-split from the uonlp/CulturaX dataset. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset. Dataset Structure Column Description text Raw Sinhala sentence source Source identifier (culturax) Split Rows train 4,707,451 Pipeline Position uonlp/CulturaX → this repo → Minuri/culturax_cleaned_version →… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-culturax.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face