CoolFace
Modelpublic

vamoruso/biscegliese-qwen2.5-gguf

sourceHugging Faceapache-2.0updated 7d agoView on Hugging Face
0likes223downloads
Model Card

🐬 Biscegliese-Qwen2.5 — Dialect LLM (GGUF)

A language model based on Qwen2.5‑7B‑Instruct, trained and optimized to understand and generate text in the Biscegliese dialect. Format: GGUF Q4_K_M, compatible with llama.cpp, LM Studio, KoboldCpp and all GGUF inference engines.


📌 Overview

Biscegliese‑Qwen2.5 is a model specialized in producing and understanding the Biscegliese dialect, with particular focus on:

  • —local idiomatic expressions
  • —dialect phonetics and morphology
  • —IT → dialect translation
  • —creative generation (poetry, dialogues, storytelling)
  • —cultural and linguistic preservation

🎯 Intended Use

Recommended Uses

  • —linguistic and dialectological research
  • —cultural and historical applications
  • —dialect content generation
  • —translation and linguistic analysis

Not Recommended

  • —medical, legal, financial domains
  • —critical or automated decision‑making
  • —sensitive or high‑risk content generation

🧠 Model Details

  • —Base model: Qwen2.5‑7B‑Instruct
  • —Format: GGUF
  • —Quantization: Q4KM
  • —Supported tasks: text‑generation, chat, translation, dialect modeling
  • —Author: Vincenzo Amoruso
  • —Version: 1.0

📚 Training Data

The model was trained on:

  • —Biscegliese dialect corpus
  • —cultural and popular texts
  • —local dialogues and transcriptions
  • —manually curated datasets

Possible Biases

  • —non‑neutral cultural representations
  • —non‑uniform phonetic variations
  • —potential errors due to limited data availability

📈 Evaluation

Qualitative Evaluation

The model was tested on:

  • —IT → dialect translations
  • —realistic dialogue generation
  • —comprehension of complex dialect sentences

Example

Input: “Come stai oggi?” Output: “Com stà ioscè?”


🔒 Safety & Ethical Considerations

The model may generate:

  • —grammatical errors
  • —inaccurate content
  • —imperfect cultural interpretations

Do not use the model for:

  • —medical diagnosis
  • —legal advice
  • —high‑risk decision‑making

⚙️ Environmental Impact

  • —Hardware: Consumer GPU RTX5090
  • —Training duration: 18 hours
  • —Estimated consumption: moderate

🛠️ Inference

Esempio con llama.cpp

bash
./llama-cli -m Qwen2.5-7B-Instruct.Q4_K_M.gguf -p "Traduci la frase in biscegliese
### Italiano:terrazzo 
### Traduzione:"