CoolFace
Datasetpublic

sh4lu-z/awesome-dataset-sinhala

Mixed Sinhala Dataset (1M+ Rows) | මිශ්‍ර සිංහල දත්ත කට්ටලය (Please find the English description below the Sinhala description) 🇬🇧 English This is a comprehensive dataset containing over one million rows of Sinhala text data. It is highly suitable for training Artificial Intelligence (AI) models and conducting Natural Language Processing (NLP) research. Dataset Details Language: Sinhala (si) Total Rows: 1,079,909 Format: Parquet (Optimized for… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/awesome-dataset-sinhala.

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
1likes95downloads
4 commits on main
077f49a7mo ago

Update README.md

sh4lu-z
29c84a47mo ago

Update README.md

sh4lu-z
64b48267mo ago

Upload dataset

sh4lu-z
49adab17mo ago

initial commit

sh4lu-z