innadark/topxgen-gemma-3-27b-and-nllb-3.3b
TopXGen: Topic-Diverse Parallel Data for Low-Resource MT Dataset Summary This dataset is a synthetic parallel dataset for 10 low-resource languages, created by applying the TopXGen pipeline with recent multilingual LLMs. It is designed for machine translation (MT) fine-tuning and few-shot experiments (as a selection pool).The pipeline works as follows: Topic-diverse paragraph generation in the target low-resource language using an LLM (generator), with… See the full description on the dataset page: https://huggingface.co/datasets/innadark/topxgen-gemma-3-27b-and-nllb-3.3b.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face