CoolFace
Modelpublic

IMISLab/Maistros-8B-Instruct

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes74downloads
Model Card

Maistros-8B-Instruct: A Greek Large Language Model adapted through Knowledge Distillation from Large Reasoning Models

‼️If the full model does not fit in your setup, you can use the official 4-bit quantized version, which uses 65% less memory.‼️

We introduce Maistros-8B-Instruct, a Greek-adapted LLM based on mistralai/Ministral-3-8B-Instruct-2512-BF16 fine-tuned using Low-Rank Adaptation (LoRA) on CulturaQA. For information regarding the model training, validation and evaluation, as well as its limitations see the arxiv preprint.

<div align="center"> <img src="Maistros-Greek.png" width="70%" alt="Maistros Greek logo"/> </div>

Model Information

  • —256k context length (approx. 150,000 Greek words).
  • —We extend the training of Ministral-3-8B-Instruct-2512-BF16 with Greek linguistic and cultural knowledge from the training part of CulturaQA.
  • —We use LoRA fine-tuning to mitigate catastrophic forgetting and retain the base models' capabilities.
  • —We merge the adapted weights from LoRA fine-tuning to the base model to produce Maistros-8B-Instruct, a specialized Greek LLM.
  • —Maistros-8B-Instruct achieves state-of-the-art performance in most Greek QA datasets, when compared to other open-weight models.

Evaluation

For the evaluation we utilize the accuracy metric for the multiple-choice datasets, while for the open-ended Cultura QA we utilize BERTScore F1%. We also utilize the instruct versions of the abbreviated models below.

DemosQAGPCRINCLUDEGreek ASEP MCQAGreek Medical MCQAPlutus QAGreek Truthful QAGreek MMLU (Greek-specific)CulturaQA
Open-Weights Models
Maistros 8B50.8364.4258.7067.2549.5473.3353.3778.1771.99
Ministral 3 8B51.6759.6254.1763.2547.9265.3352.5176.2371.03
Krikri 8B49.5054.8150.5463.0845.3764.4454.8371.0471.31
Plutus 8B45.6750.0048.3762.9239.3557.3334.5270.3867.44
EuroLLM v2 9B41.5053.8539.1346.0831.7142.6736.7258.1770.33
Gemma 3n E4B47.1760.1050.0057.7543.7553.7846.7671.3969.10
Qwen 3 8B48.8331.7349.2854.5836.6463.5642.7267.5768.73
Proprietary Models
Gemini 3 flash55.6788.4688.7794.7592.8289.7888.6295.0373.97
GPT-5 mini53.0077.4074.4678.9278.0176.8975.8987.4975.09

How to load and run the model.

Use the following code to run the model locally or you can host the model using vLLM.

python
from transformers import AutoTokenizer, Mistral3ForConditionalGeneration, set_seed

# Set the model path, device and a random seed for reproducibility.
model_path = 'IMISLab/Maistros-8B-Instruct'
device = 'cuda'
set_seed(42)

# Loading the model tokenizer.
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code = True)

# Causal Language Models predict tokens from left to right and use EOS token for padding.
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = 'right'

# Load the model from the path to the device and set it in evaluation mode.
model = Mistral3ForConditionalGeneration.from_pretrained(model_path, device_map = device, trust_remote_code = True)
model.eval()

# Set the system, instruction and user prompts.
system_prompt = 'Είσαι ο Μαΐστρος, ένα εξαιρετικά ανεπτυγμένο μοντέλο Τεχνητής Νοημοσύνης για την Ελληνική γλώσσα.\nΈχεις δημιουργηθεί απο το IMIS Lab του Πανεπιστημιού Πατρών.'
instruction_prompt = 'Παρακαλώ απάντησε στην παρακάτω ερώτηση.'
user_prompt = 'Τι είναι η Ακρόπολη των Αθηνών;'

# Defining the message template.
messages = [
    {'role': 'system', 'content': [{'type': 'text', 'text': system_prompt}]},
    {'role': 'user', 'content': [{'type': 'text', 'text': '\n\n'.join((instruction_prompt, user_prompt))}]}
]

# Applying the tokenizer chat template.
tokenized = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt = True,  
    return_tensors = 'pt', 
    return_dict = True
)

# Sending the tokenized instances to the device.  
tokenized = {k: v.to(device) for k, v in tokenized.items()}
input_len = len(tokenized['input_ids'][0])

# Generating the model output.
output = model.generate(
    **tokenized,
    max_new_tokens = 1024,
    do_sample = False, # Equivalent to temperature = 0.0
    temperature = None,
    top_p = None,
    top_k = None
)

# Decoding the assistant part of the output and printing it.
decoded_output = tokenizer.decode(output[0][input_len:], skip_special_tokens = True)
print(decoded_output)

Contact

If you have any questions/feedback about the dataset please e-mail one of the following authors:

giarelis@ceid.upatras.gr
cmastrokostas@ac.upatras.gr
karacap@upatras.gr

Citation

@misc{
  giarelis2026maistrosgreeklargelanguage,
  title = {Maistros: A Greek Large Language Model Adapted Through Knowledge Distillation From Large Reasoning Models}, 
  author = {Nikolaos Giarelis and Charalampos Mastrokostas and Nikos Karacapilidis},
  year = {2026},
  eprint = {2605.01870},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url = {https://arxiv.org/abs/2605.01870}, 
}