CoolFace
Datasetpublic

lubzo/marathi-alpaca-cleaned-translated

Marathi Alpaca Cleaned Translated A Marathi translation of the 51,760-row Alpaca-Cleaned instruction-tuning dataset — Unsloth's hosted fork of yahma/alpaca-cleaned, which fixes hallucinations, empty outputs, and formatting errors found in the original Stanford Alpaca-52k dataset. Translated using Meta's facebook/nllb-200-distilled-600M model. Built to reproduce and evaluate the Marathi instruction-tuning experiment from Khade et al., CHiPSAL 2025. The original paper translated… See the full description on the dataset page: https://huggingface.co/datasets/lubzo/marathi-alpaca-cleaned-translated.

sourceHugging Facecc-by-nc-4.0updated 7d agoView on Hugging Face
0likes72downloads

lubzo/marathi-alpaca-cleaned-translated · main · files are served by the source, never re-hosted here