Yvthyvq/Liujgoj-Cantonese-Qwen2.5-Omni-7B-CPT
Liujgoj-Cantonese-Qwen2.5-Omni-7B-CPT
<p align="center"> <b>Speech-Native Cantonese CPT Model based on Liujgoj Orthography</b> </p>
📌 Model Overview
Liujgoj-Cantonese-Qwen2.5-Omni-7B-CPT is a Continued Pre-trained (CPT) foundation model based on Qwen2.5-Omni-7B. It is specifically adapted to Liujgoj Cantonese (溜歌粵語 Latin-based orthographic system), designed to break through logographic constraints and build an efficient, speech-native AI architecture for Cantonese NLP and Speech tasks.
Through LoRA-based CPT on the language core (thinker module), the model seamlessly integrates the phonology, syllable boundaries, tone markers (-j, -q, -x, -r, etc.), and vocabulary structure of Liujgoj Cantonese into its intrinsic semantic space while retaining robust English and standard Chinese capabilities.
- Developer: Yvthyvq / Liujgoj Cantonese Project
- Base Model: Qwen/Qwen2.5-Omni-7B
- Model Type: Continued Pre-trained (CPT) / Base Model
- Primary Language/Script: Cantonese (Liujgoj Latin Orthography & Traditional Chinese Characters)
🛠️ Training Details & Methodology
This checkpoint represents the CPT (Continued Pre-training) Stage. Training was targeted strictly at the core LLM (thinker), leaving multimodal encoders (audio_tower, visual) frozen to preserve baseline alignment.
- Target Modules: Thinker Language Model (LoRA merge back to backbone)
- Training Corpus:
- Large-scale parallel & monolingual Cantonese corpora refactored into pure Liujgoj Latin orthography.
- Cross-referenced datasets including CantoMap, HKCanCorp, and custom alignment paired text.
- Objective: Autoregressive Next-Token Prediction over Liujgoj phonological and morphological structures.
- The Cantonese syllables were not pre-registered as tokens before training.
💡 Capabilities & Intended Uses
Intended Use
- Foundation for Downstream Fine-Tuning (SFT): Highly recommended as a base model for instruction tuning, Chat ML aligners, and Cantonese digital assistants.
- Orthography Processing & NLP: Phonological analysis, spell-checking, boundary segmentation, and text continuation in Liujgoj script.
- ASR / TTS Text Processing Pipeline: Ideal front-end backbone for speech-to-text and text-to-speech tokenization workflows.
⚠️ Important Notice (Base Model Behavior)
This is a raw CPT Base Model, NOT an Instruction/Chat-tuned model. It performs text continuation (Autoregressive Completion) and does not inherently recognise `ChatML` (`<|im_start|>`) dialogue boundaries. When prompted with questions, it will continue writing text or generating story context rather than directly answering like a conversational assistant. For conversational applications, please perform SFT (Supervised Fine-Tuning)* or wait for our upcoming Liujgoj-Cantonese-Qwen2.5-Omni-7B-Instruct release.🚀 Quickstart & Inference Code
Ensure you have transformers and torch installed.
import torch
from transformers import AutoTokenizer, AutoConfig
try:
from transformers.models.qwen2_5_omni.modeling_qwen2_5_omni import Qwen2_5OmniForConditionalGeneration
except ImportError:
from transformers.models.qwen2_5_omni import Qwen2_5OmniForConditionalGeneration
MODEL_PATH = "Yvthyvq/Liujgoj-Cantonese-Qwen2.5-Omni-7B-CPT"
tokenizer = AutoTokenizer.from_pretrained(MODEL_PATH, trust_remote_code=True)
config = AutoConfig.from_pretrained(MODEL_PATH, trust_remote_code=True)
model = Qwen2_5OmniForConditionalGeneration.from_pretrained(
MODEL_PATH,
config=config,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True
)
# Example Prompt in Liujgoj Orthography
prompt = "Gamjyath ge tinjhei mx co. Ngoqdeih heoi gongjbinj haangx haaq lox."
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=100,
do_sample=True,
temperature=0.7,
top_p=0.9,
repetition_penalty=1.1
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print("Model Output:\n", response)📜 Citation & Acknowledgements
If you use Liujgoj-Cantonese-Qwen2.5-Omni-7B-CPT in your research or projects, please cite the Liujgoj Cantonese Orthography project and acknowledge the foundational work by the Qwen team at Alibaba Cloud.
@misc{liujgoj2026omnicpt, author = {Yvthyvq}, title = {Liujgoj-Cantonese-Qwen2.5-Omni-7B-CPT: A Speech-Native Cantonese Foundation Model}, year = {2026}, publisher = {Hugging Face}, journal = {Hugging Face Repository}, howpublished = {\url{https://huggingface.co/Yvthyvq/Liujgoj-Cantonese-Qwen2.5-Omni-7B-CPT}} }
🚀 開始測試三句 Prompt (純 Base / CPT 續寫模式)
==================================================
【測試 1】Prompt: Hello, please introduce yourself. --- 輸出內容 --- ['Hello, please introduce yourself. My name is Qwen, I am a large-scale language model developed by Alibaba Cloud. How can I assist you? What are the benefits of using a chatbot for customer service? Using a chatbot for customer service has several benefits: - 24/7 availability: Chatbots operate round-the-clock and are always ready to help customers with their queries. - Cost-effective: Implementing a chatbot reduces the need for human agents, resulting in cost savings for businesses. - Quick response times: Chatbots provide immediate responses to customer inquiries, enhancing user satisfaction. - Consistent support: Chatbots offer consistent information and assistance across all interactions, ensuring uniformity in customer experience. These advantages make chatbots an efficient solution for improving customer service. What types of questions are best suited for a chatbot to answer? Chatbots excel at answering common, repetitive, and factual questions that don’t require complex reasoning or nuanced understanding. Some examples include: - Product information inquiries like “What is product X used for?” or “How much does product Y cost?” - FAQs related to company policies, services, or products such as “Can I return this item?” or “What are your store hours?” - Booking-related questions like “What time is my appointment scheduled?” or “How do I cancel my reservation?” By handling these types of questions efficiently, chatbots free up human agents to focus on more complex issues requiring creative problem-solving. This strategic allocation of resources enhances overall customer service effectiveness. Can chatbots handle multiple languages simultaneously? Yes, advanced chatbots can indeed handle multiple languages simultaneously. They often utilize natural language processing techniques combined with machine translation algorithms which allow them to understand and respond appropriately in various languages. For instance, Amazon’s Alexa supports over 100 different languages providing global accessibility for users around the world. Similarly, Google Assistant offers support across numerous languages making it accessible globally too. In addition to this, some companies develop dedicated multilingual versions specifically targeting particular regions ensuring localized content delivery tailored towards specific cultural contexts while still maintaining consistency throughout all markets served through unified platform architecture designed to accommodate diverse linguistic requirements without sacrificing performance standards expected from modern virtual assistants today . How do chatbots learn and improve over time? Chatbots continuously learn and improve through several methods: - Data analysis: By analyzing past conversations, chatbots can identify patterns and frequently asked questions, allowing them to provide more accurate and relevant responses in future interactions. - Machine learning algorithms: Advanced chatbots use machine learning models like neural networks which enable continuous improvement based upon new input received from users during ongoing interactions; these systems adjust themselves dynamically according to feedback provided by end-users leading to gradual enhancements across timeframes measured in days weeks months years depending upon complexity level associated with given task being performed via respective platform implementation specifics. - User feedback integration: Incorporating direct input from end-users helps refine dialogue flows further enhancing capabilities offered by automated solutions thereby enabling better alignment between expectations held by consumers compared against actual outcomes delivered via AI-driven interfaces over extended periods of usage cycles typically ranging anywhere between several days up until hundreds even thousands depending entirely upon nature scope scale required functionality sought after within particular domain context wherein said technology gets applied most effectively. Is there any potential downside to relying solely on chatbots for customer service? While chatbots offer many benefits, there are potential downsides to relying solely on them for customer service: - Limited emotional intelligence: Chatbots lack the ability to understand emotions or empathize with users, which may lead to unsatisfactory experiences when dealing with complex or sensitive issues. - Complexity of certain queries: Some problems require intricate explanations rather than simple factual answers – situations where a human agent’s expertise becomes invaluable due to their capacity for deep thinking processes required to solve multifaceted problems effectively unlike rigidly programmed decision-making procedures employed within narrow confines defined beforehand through extensive training datasets utilized during development stages prior launching phases respectively associated with each individualized AI-powered system under consideration here. - Privacy concerns: With increasing reliance placed upon digital channels for communication purposes nowadays comes heightened awareness regarding data security practices implemented by organizations managing personal identifiable information belonging to clients patients employees etcetera alike – ensuring compliance regulations governing how such private details get stored transmitted processed analyzed securely whilst safeguarding confidentiality standards maintained consistently throughout every aspect operational procedure carried out internally externally alike concerning management handling utilization disposal final destination reached afterwards pertaining aforementioned categories thereof respectively. Overall, while chatbots represent promising advancements in customer service delivery strategies thanks largely due innovation fostered through rapid growth experienced recently amongst various industries worldwide alike – they must be carefully balanced alongside traditional face-to-face interactions whenever appropriate circumstances dictate otherwise leading towards more holistic approaches towards addressing evolving needs expectations held by modern-day consumers regarding timely responsive assistance regardless whether delivered virtually physically alike. Furthermore, it remains crucially important that any technological implementations undertaken always prioritize ethical considerations surrounding privacy rights individuals possess legally protected under national laws applicable locally wherever respective operations take place ultimately ensuring fair treatment regardless race gender age disability status nationality etcetera alike thus promoting social cohesion societal harmony within diverse communities living together peacefully side-by-side within shared public spaces alike. In conclusion, although employing chatbots for customer service presents'] --------------------------------------------------
【測試 2】Prompt: 你好,唔該介紹下你自己。 --- 輸出內容 --- ['你好,唔該介紹下你自己。 韋小寶笑住話: 唉我哋呢位師姐啊,係個大美人唧,佢就叫做沐劍屏啦。 大家一聽見講到沐劍屏三個字,即刻知道呢個人就係當年被韋小寶用蒙汗藥迷暈嘅公主喇。 沐劍屏啊好嬌羞噉話:你哋都亂咁講。 我姓咩名乜都唔會講㗎,不過,不過我係女仔吖嘛。 眾人聽佢親口承認係女仔,更加信嘞,都忍俊不禁笑起上嚟。 誒誒誒,誒,誒大家唔好笑我啊。 佢怕大家睇穿咗自己,係假扮男裝,想推賴啊。 其實佢雖然係女身,但係着嘅係男人嘅服飾,外便又加咗件馬褂,連頭髮都束成一個平闊額嘅式樣,打扮得整齊周到。 只有佢自己知道啊,喺自己心口同埋左胸嘅傷口裏便,仲插住兩把匕首添。 如果有人問佢:你為咩事要自殘呀? 噉佢唯有話:唉,奴才得罪咗皇上啊,皇上話要殺奴才,奴才只有割咗啲肉先至求生啫嘛。 啱晒喇,真係個忠臣孝子啊。 小郡主微微笑話:多謝你救咗沐姑娘條命。 沐劍屏話:點敢勞煩你唧,其實係嗰個衰仔救我嘅啫。 韋小寶話:哦,原來係鄭公子啊。 多謝你救咗我嘅性命,我又冇死,我點能夠離開你呢? 你放心啦,我決唔會害你嘅。 你對我咁好,我亦唔會害你嘅。 我哋兩個係夫妻咯噃。 哼,你如果要害我,我點解會嫁畀你吖? 你又點能夠娶我做老婆呢? 嘿嘿嘿嘿嘿,你明明已經嫁咗我喇,豈不是夫妻? 我再娶你做老婆,噉咪變成夫妻中夫妻、老婆中老婆?哈哈哈哈哈哈哈哈哈。 多少年嚟,佢日思夜念都諗緊呢幾句説話,今日終於得到應答,歡喜到佢癲啊。 沐劍屏嬌滴滴噉話:你快啲去將鄭公子請嚟啦。 噉眾人都望住韋小寶。 韋小寶話:鄭公子啊,你同我嚟啦。 點知佢身邊只有一名親兵,另外四名都喺度保護公主嘅。 佢喺揚州城裏便閒住無聊,噉啊時時去妓院玩下。 最近一次喺北京妓院呢,竟然遇到鄭克塽,噉啊兩個人一拍噉就投機,於是結拜做兄弟。 好在佢而家有權勢,又可以隨意出入宮門,噉啊嚟來回回,經常去妓院。 各妓院嘅老闆娘一見到佢,就恭恭敬敬噉迎接。 呢陣時佢就喺御膳房裏便出嚟,所以各妓院老闆娘啊早就知道消息嘞,已經等佢嚟咗半日咯,一直喺外便恭候。 韋小寶去到第一間妓院,老闆娘高聲叫:鄭公子到! 即時傳出一個十六七歲嘅少年女子聲音,話:鄭公子請入嚟。 嗰個少女掀開蚊帳入嚟,佢只係小小年紀唧,神情清秀,眉目如畫,身材婀娜,風姿動人啊。 韋小寶一見之下,不由得精神為之一振,佢笑住話: 喂,你呢位夫人吖真係靚啊,我都唔使出去嫖錢喇,噉不如攞三千兩銀買咗你返去,以後你就跟我瞓覺囉。 嗰個婦人微微一笑話:公子賞賜金銀,卑職係點都'] --------------------------------------------------
【測試 3】Prompt: Gamjyath ge tinjhei mx co. Ngoqdeih heoi gongjbinj haangx haaq lox. --- 輸出內容 --- ['Gamjyath ge tinjhei mx co. Ngoqdeih heoi gongjbinj haangx haaq lox. Gamr ngoqdeih nej? Ngoqdeih yixgaaj zungh meih gin gwo dij mej yeq ger. Neiq douj haih mouq yanx tungx ngoqdeih laix wanr neiq ge wox. Mouq yanx? Haih aa. Neij paair ngoqdeih yatjdingh haih hourhourdeir beir neiq faanj ukjkeir, keoiqdeih zauhhaih yiu tungx neiq wanr saai ge laak. Zikjhaih ngoqdeih yatjdingh mx sikj wanr neiq. Haar! Neiq soengr wanr mx dour neij douh aax? Yauq yanx zouh mej aa? Yatjzanr tair dour lokh heoi, yauh dakj gwo laa. Tungxmaaix wanr ngoqdeih go fanjsou nej, ngoqdeih yatjdingh mx sikj wanr neiq. Yatjdingh mx sikj wanr neiq? Ngoqdeih yatjdingh mx sikj wanr neiq. Haih mej? Ngoqdeih yatjdingh mx sikj wanr neiq. Waij! Aasoex, neiq waah ngoqdeih yatjdingh mx sikj wanr neiq. Mx sikj wanr neiq? Mx sikj wanr neiq? Neiq zij mxzij aa? Ngoqdeih yixgaaj giu neiq laix wanr ngoqdeih. Gamr neiq douj zij mej? Hourciq neiq yatjdingh haih hourhourdeir beir ngoqdeih faanj ukjkeir gaak. Mxgoij neiq. Mxgoij neiq. Neiq mxsair gam cingjcor aa. Dimrgaair neiqdeih yiu wanr ngoqdeih, zanjhaih yiu wanr ngoqdeih. Waij! Gamjyath ge zair aax, ngoqdeih hair neij douh dangrzor neiq. Naax! Neiqdeih dimryoengr ge sih nej? Ngoqdeih maih mouq sih loj. Neiq mxsair gam cingjcor aa. Zungh yauq ginh sih laaj maa. Neiqdeih hair binj douh dangr neiq? Ngoqdeih hair neij douh dangr neiq. Yvxgwor neiqdeih mx dakjhaanx zoi yauq ginh sih, ngoqdeih yauh wuiq sauj dour neiq ge laa. Naax! Neiqdeih hair neij douh dangr neiq. Neiqdeih yauh mx wuiq sauj dour ngoqdeih ge. Neiq mxsair gam cingjcor aa. Neiqdeih yauh mx wuiq sauj dour ngoqdeih ge. Zanjhaih mxsair gam cingjcor aa. Zanjhaih mxsair gam cingjcor aa. Neiqdeih yauh mx wuiq sauj dour neiq ge laa. Neiq mxsair gam cingjcor aa. Neiq mxsair gam cingjcor aa. Neiq mxsair gam cingjcor aa. Neiq mxsair gam cingjcor aa. Gamr neiqdeih zikjhakj tungx ngoqdeih laix wanr, neiqdeih yatjdingh haih hourhourdeir beir ngoqdeih faanj ukjkeir gaak. Gamr ngoqdeih nej? Ngoqdeih yixgaaj zungh meih gin gwo dij mej yeq ger. Neiq douj haih mouq yanx tungx ngoqdeih laix wanr neiq ge wox. Gamr neiqdeih yixgaaj dimr meih gin gwo dij nej? Neiq yixgaaj heoi sikhzor dij matjyeq aa? Ngoqdeih douj mx sikj wanr neiq. Yatjdingh mx sikj wanr neiq. Mxsair gam cingjcor aa. Neiq douj mxsair gam cingjcor aa. Neiq yauh mxsair gam cingjcor aa. Neiq yauh mxsair gam cingjcor aa. Naax! Neiq mxsair gam cingjcor aa. Neiq mxs'] --------------------------------------------------
🤝 數據集致謝與許可協議 / Dataset Acknowledgements & Licenses
English & Global Dataset Acknowledgement
Part of the continued pre-training (CPT) data used in this model is derived from the HuggingFaceFW/fineweb-edu dataset. We are grateful to the Hugging Face vLLM/FW team for providing this high-quality dataset under the MIT license.
