Lo-Renz-O/malagasy-sentence
Overview This dataset consists of clean, structured sentences extracted via Optical Character Recognition (OCR) from approximately 1GB of Malagasy thesis documents. These documents were collected based on educational, cultural, and linguistic themes. The dataset is saved in CSV format, and is particularly useful for NLP tasks involving sentence-level modeling in Malagasy — a low-resource language. Dataset Details Language: Malagasy Source: OCR'd academic thesis… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/malagasy-sentence.
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload README.md with huggingface_hub
Upload dataset
initial commit
