CoolFace
Datasetpublic

Navanjana/sinhala-articles

Sinhala Articles Dataset A large-scale, high-quality Sinhala text corpus curated from diverse sources including news articles, Wikipedia entries, and general web content. This dataset is designed to support a wide range of Sinhala Natural Language Processing (NLP) tasks. 📊 Dataset Overview Name: Navanjana/sinhala-articles Total Samples: 2,148,688 Languages: Sinhala (si) Features: text: A single column containing Sinhala text passages. Size: Approximately 1M… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/sinhala-articles.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
1likes18downloads
17 commits on main
505cb2c1y ago

Update README.md

Navanjana
2fbf4c71y ago

Update README.md

Navanjana
911a7871y ago

Update README.md

Navanjana
423a7001y ago

Upload sinhala_dataset.csv with huggingface_hub

Navanjana
74cc8b81y ago

Update README.md

Navanjana
9095fcb1y ago

Update README.md

Navanjana
4a8a6861y ago

Update README.md

Navanjana
f2745d71y ago

Update README.md

Navanjana
544729c1y ago

Upload sinhala_dataset_new.csv with huggingface_hub

Navanjana
a7bc83d1y ago

Update README.md

Navanjana
8b215001y ago

Update README.md

Navanjana
98e6fc81y ago

Update README.md

Navanjana
1246a8f1y ago

Update README.md

Navanjana
29194411y ago

Update README.md

Navanjana
89f795d1y ago

Upload new_dataset.csv with huggingface_hub

Navanjana
f2d41891y ago

Upload sinhala_dataset.csv with huggingface_hub

Navanjana
00249841y ago

initial commit

Navanjana