CoolFace
Modelpublic

pguerrero-igutierrez/Latxa-Qwen3-8B-Literary-v1-tokenmatched-ca-eu

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes7downloads
Model Card

Latxa-Qwen3-8B-Literary-v1-tokenmatched-ca-eu

LoRA adapter for Catalan-to-Basque literary machine translation, fine-tuned from `HiTZ/Latxa-Qwen3-VL-8B-Instruct`.

This model is part of the MT Domain Adaptation CA-EU collection and was trained for experiments on domain adaptation under low-resource Catalan-Basque translation conditions.

Model details

  • —Model type: PEFT LoRA adapter for causal language modeling
  • —Base model: HiTZ/Latxa-Qwen3-VL-8B-Instruct
  • —Task: Catalan → Basque machine translation
  • —Domain: Literary text
  • —Languages: Catalan (ca), Basque (eu)
  • —Adapter method: LoRA
  • —Training setup: token-budget matched literary training data, designed to match the clinical-domain token budget for fairer cross-domain comparison

Intended use

This adapter is intended for research on Catalan-Basque machine translation and domain adaptation, especially for literary-domain translation.

Example instruction format:

text
Tradueix aquest text del català al basc:

{source_text}

How to use

python
from transformers import AutoTokenizer, Qwen3VLForConditionalGeneration
from peft import PeftModel

base_model = "HiTZ/Latxa-Qwen3-VL-8B-Instruct"
adapter = "pguerrero-igutierrez/Latxa-Qwen3-8B-Literary-v1-tokenmatched-ca-eu"

tokenizer = AutoTokenizer.from_pretrained(adapter, trust_remote_code=True)

model = Qwen3VLForConditionalGeneration.from_pretrained(
    base_model,
    device_map="auto",
    torch_dtype="auto",
    trust_remote_code=True,
)

model = PeftModel.from_pretrained(model, adapter)
model.eval()

source = "La bellesa és com l'alcohol o el confort, et acostumes a ella i deixes de prestar-li atenció."
prompt = f"Tradueix aquest text del català al basc:\n\n{source}"

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=256,
    do_sample=False,
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Training data

The adapter was fine-tuned on Catalan-to-Basque literary translation pairs derived from a literary backtranslated corpus.

The training split was selected using a token-budget matching strategy: the amount of literary-domain training and validation data was sampled to match the source-token budget of the corresponding clinical-domain setup. This was done to support fairer comparisons between domain-adapted models.

Training procedure

The model was trained with LoRA using the following configuration:

ParameterValue
Base modelHiTZ/Latxa-Qwen3-VL-8B-Instruct
Epochs3
Max sequence length768
Per-device batch size4
Gradient accumulation8
Learning rate5e-5
LR schedulercosine
Warmup ratio0.05
Seed42
LoRA rank16
LoRA alpha32
LoRA dropout0.05
Target modulesq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Quantization during training4-bit NF4 when CUDA was available
Best checkpoint step300
Best validation BLEU4.175

Evaluation

Evaluation was performed on a held-out Catalan-to-Basque literary test set.

DirectionchrF++BLEUTERCOMETLength ratioSamples
ca → eu24.732.2897.0365.410.7724,261

Out-of-scope use

This model is not intended for legal, medical, safety-critical, or fully automated professional translation workflows without human review.

Citation

If you use this adapter, please cite or acknowledge the MT Domain Adaptation CA-EU project and the base model:

bibtex
@misc{latxa_qwen3_literary_tokenmatched_ca_eu,
  title = {Latxa-Qwen3-8B-Literary-v1-tokenmatched-ca-eu},
  author = {Gutierrez Fandiño, Iker and Guerrero Castelló, Paula},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/pguerrero-igutierrez/Latxa-Qwen3-8B-Literary-v1-tokenmatched-ca-eu}}
}

Framework versions

  • —PEFT: 0.17.1
  • —Transformers-compatible adapter