CoolFace
Datasetpublic

innadark/topxgen-gemma-3-27b-and-nllb-3.3b

TopXGen: Topic-Diverse Parallel Data for Low-Resource MT Dataset Summary This dataset is a synthetic parallel dataset for 10 low-resource languages, created by applying the TopXGen pipeline with recent multilingual LLMs. It is designed for machine translation (MT) fine-tuning and few-shot experiments (as a selection pool).The pipeline works as follows: Topic-diverse paragraph generation in the target low-resource language using an LLM (generator), with… See the full description on the dataset page: https://huggingface.co/datasets/innadark/topxgen-gemma-3-27b-and-nllb-3.3b.

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes41downloads
1 commits on main
b15edff9mo ago

Duplicate from almanach/topxgen-gemma-3-27b-and-nllb-3.3b

innadark, ArmelR