pguerrero-igutierrez/Latxa-Qwen3-8B-Literary-v1-tokenmatched-ca-eu
Latxa-Qwen3-8B-Literary-v1-tokenmatched-ca-eu
LoRA adapter for Catalan-to-Basque literary machine translation, fine-tuned from `HiTZ/Latxa-Qwen3-VL-8B-Instruct`.
This model is part of the MT Domain Adaptation CA-EU collection and was trained for experiments on domain adaptation under low-resource Catalan-Basque translation conditions.
Model details
- Model type: PEFT LoRA adapter for causal language modeling
- Base model:
HiTZ/Latxa-Qwen3-VL-8B-Instruct - Task: Catalan → Basque machine translation
- Domain: Literary text
- Languages: Catalan (
ca), Basque (eu) - Adapter method: LoRA
- Training setup: token-budget matched literary training data, designed to match the clinical-domain token budget for fairer cross-domain comparison
Intended use
This adapter is intended for research on Catalan-Basque machine translation and domain adaptation, especially for literary-domain translation.
Example instruction format:
Tradueix aquest text del català al basc:
{source_text}How to use
from transformers import AutoTokenizer, Qwen3VLForConditionalGeneration
from peft import PeftModel
base_model = "HiTZ/Latxa-Qwen3-VL-8B-Instruct"
adapter = "pguerrero-igutierrez/Latxa-Qwen3-8B-Literary-v1-tokenmatched-ca-eu"
tokenizer = AutoTokenizer.from_pretrained(adapter, trust_remote_code=True)
model = Qwen3VLForConditionalGeneration.from_pretrained(
base_model,
device_map="auto",
torch_dtype="auto",
trust_remote_code=True,
)
model = PeftModel.from_pretrained(model, adapter)
model.eval()
source = "La bellesa és com l'alcohol o el confort, et acostumes a ella i deixes de prestar-li atenció."
prompt = f"Tradueix aquest text del català al basc:\n\n{source}"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))Training data
The adapter was fine-tuned on Catalan-to-Basque literary translation pairs derived from a literary backtranslated corpus.
The training split was selected using a token-budget matching strategy: the amount of literary-domain training and validation data was sampled to match the source-token budget of the corresponding clinical-domain setup. This was done to support fairer comparisons between domain-adapted models.
Training procedure
The model was trained with LoRA using the following configuration:
Evaluation
Evaluation was performed on a held-out Catalan-to-Basque literary test set.
Out-of-scope use
This model is not intended for legal, medical, safety-critical, or fully automated professional translation workflows without human review.
Citation
If you use this adapter, please cite or acknowledge the MT Domain Adaptation CA-EU project and the base model:
@misc{latxa_qwen3_literary_tokenmatched_ca_eu,
title = {Latxa-Qwen3-8B-Literary-v1-tokenmatched-ca-eu},
author = {Gutierrez Fandiño, Iker and Guerrero Castelló, Paula},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/pguerrero-igutierrez/Latxa-Qwen3-8B-Literary-v1-tokenmatched-ca-eu}}
}Framework versions
- PEFT: 0.17.1
- Transformers-compatible adapter
