sh4lu-z/awesome-dataset-sinhala
Mixed Sinhala Dataset (1M+ Rows) | මිශ්ර සිංහල දත්ත කට්ටලය (Please find the English description below the Sinhala description) 🇬🇧 English This is a comprehensive dataset containing over one million rows of Sinhala text data. It is highly suitable for training Artificial Intelligence (AI) models and conducting Natural Language Processing (NLP) research. Dataset Details Language: Sinhala (si) Total Rows: 1,079,909 Format: Parquet (Optimized for… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/awesome-dataset-sinhala.
195
Update README.md
Update README.md
Upload dataset
initial commit
