tartuNLP/pale-madlad-data
license: mit PaLe-MADLAD Data Data used for training the PaLe-MADLAD model to translate from Proper Karelian, Livvi, Ludian, and Veps to Russian and vice versa. Every dataset entry represents a single text and comes as a list of sentences supplemented (where possible) with a list of translations into Russian. Our sources include: VepKar: various articles, Biblical texts, folklore, and more in Proper Karelian, Livvi, Ludian, and Veps, mostly translated into… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/pale-madlad-data.
Update README.md
Added dataset info
Update README.md
Update README.md
Randomized data.
Fixed vepkar data.
Swapped entries in README.
Added dataset files.
Added config to readme.
initial commit
