innadark/topxgen-gemma-3-27b-and-nllb-3.3b
TopXGen: Topic-Diverse Parallel Data for Low-Resource MT Dataset Summary This dataset is a synthetic parallel dataset for 10 low-resource languages, created by applying the TopXGen pipeline with recent multilingual LLMs. It is designed for machine translation (MT) fine-tuning and few-shot experiments (as a selection pool).The pipeline works as follows: Topic-diverse paragraph generation in the target low-resource language using an LLM (generator), with… See the full description on the dataset page: https://huggingface.co/datasets/innadark/topxgen-gemma-3-27b-and-nllb-3.3b.
041
Duplicate from almanach/topxgen-gemma-3-27b-and-nllb-3.3b
