CoolFace
Datasetpublic

Lo-Renz-O/malagasy-sentence

Overview This dataset consists of clean, structured sentences extracted via Optical Character Recognition (OCR) from approximately 1GB of Malagasy thesis documents. These documents were collected based on educational, cultural, and linguistic themes. The dataset is saved in CSV format, and is particularly useful for NLP tasks involving sentence-level modeling in Malagasy — a low-resource language. Dataset Details Language: Malagasy Source: OCR'd academic thesis… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/malagasy-sentence.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes26downloads
9 commits on main
dbbe9161y ago

Update README.md

Lo-Renz-O
2a8734c1y ago

Update README.md

Lo-Renz-O
5811e071y ago

Update README.md

Lo-Renz-O
14231961y ago

Update README.md

Lo-Renz-O
51ddb3c1y ago

Update README.md

Lo-Renz-O
e3440791y ago

Update README.md

Lo-Renz-O
886a7751y ago

Upload README.md with huggingface_hub

Lo-Renz-O
3a13ba81y ago

Upload dataset

Lo-Renz-O
326bfcf1y ago

initial commit

Lo-Renz-O