Navanjana/sinhala-articles
Sinhala Articles Dataset A large-scale, high-quality Sinhala text corpus curated from diverse sources including news articles, Wikipedia entries, and general web content. This dataset is designed to support a wide range of Sinhala Natural Language Processing (NLP) tasks. 📊 Dataset Overview Name: Navanjana/sinhala-articles Total Samples: 2,148,688 Languages: Sinhala (si) Features: text: A single column containing Sinhala text passages. Size: Approximately 1M… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/sinhala-articles.
Update README.md
Update README.md
Update README.md
Upload sinhala_dataset.csv with huggingface_hub
Update README.md
Update README.md
Update README.md
Update README.md
Upload sinhala_dataset_new.csv with huggingface_hub
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload new_dataset.csv with huggingface_hub
Upload sinhala_dataset.csv with huggingface_hub
initial commit
