CoolFace
Modelpublic

Yvthyvq/Liujgoj-Cantonese-Qwen2.5-Omni-7B-ASR

sourceHugging Faceapache-2.0updated 2h agoView on Hugging Face
0likes64downloads
Model Card

大規模測試:

對 CanCLID/zoengjyutgaai 中mouzaakdung音頻(https://huggingface.co/datasets/CanCLID/zoengjyutgaai/tree/main/source/mouzaakdung)嘅轉錄結果:

(MouxZaakhdungj_omni.xlsx)

Whisper vs Omni ASR 轉寫結果比較

比較指標Whisper 轉寫結果Omni 轉寫結果
總轉寫樣本數4,742 條4,742 條
平均文本長度98.93 字元102.50 字元
平均文本相似度71.25%71.25%
混入漢字樣本數827 條 (17.44%)10 條 (0.21%)
漢字總字數16,941 字57 字
重複/死循環現象174 條 (3.67%)1 條 (0.02%)
大寫字母總數2,388 個8,656 個

📊 Benchmark & Evaluation Report

我們針對 4 個 ASR 模型在 Liujgoj 粵語羅馬字聽寫上的表現進行了深度評測。詳細報告與數據飛輪路線圖已更新至倉庫:

  • —📖 [線上閱讀完整 Markdown 評測報告](./BENCHMARK.md)
  • —📥 [下載高清晰 PDF 評測報告](https://huggingface.co/Yvthyvq/Liujgoj-Cantonese-Qwen2.5-Omni-7B-ASR/resolve/main/BENCHMARK.pdf)

喺使用本ASR對數據集(https://huggingface.co/datasets/CanCLID/zoengjyutgaai/tree/main/source/mouzaakdung)進行轉寫時,有1句(約佔總數據 0.02%)出現完全相同的 yiu 無限重複嘅語句。

0290148.wav:毛澤東嗰張闊大嘅木床,而家顯得係咁平整、光滑、潔淨。 { "file": "0290148.wav", "path": "/root/autodl-tmp/mouxzaakhdungj/029_0148.wav", "transcription": "Mouqzauh Dungjgwok gojzungr fuhdaaih ge mukhcongr, yixgaaj yauh yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu yiu y" },

原因排查:

1,訓練數據有三句含有“yiu yiu”嘅句子,擴增爲20句,導致出現吸附子; 2,有可能是 Softmax 計算時撞正某個浮點數臨界值。


Liujgoj-Cantonese-Qwen2.5-Omni-7B-ASR

A Cantonese Automatic Speech Recognition model that directly transcribes spoken Cantonese into Liujgoj Cantonese Orthography (溜歌粵語羅馬字).

Overview

Liujgoj-Cantonese-Qwen2.5-Omni-7B-ASR is a specialized Automatic Speech Recognition (ASR) model built through a three-stage training pipeline:

text
Qwen2.5-Omni-7B
        │
        ▼
Liujgoj Cantonese CPT
        │
        ▼
Liujgoj Cantonese SFT
        │
        ▼
Liujgoj Cantonese ASR

Starting from Qwen2.5-Omni-7B, the model was first adapted to the Liujgoj Cantonese language space through Continued Pretraining (CPT), then further structured through Supervised Fine-Tuning (SFT), and finally fine-tuned for Automatic Speech Recognition.

The final model is designed to perform:

Spoken Cantonese Audio → Liujgoj Cantonese

Unlike conventional Cantonese ASR systems that typically transcribe speech into Chinese characters, this model directly generates Liujgoj Romanized Cantonese, including syllable structure, tone letters, word boundaries, and contextual linguistic information.


Training Pipeline

Stage 1 — Continued Pretraining (CPT)

The base Qwen2.5-Omni-7B model was continued-pretrained on approximately 14M of Liujgoj-related training data.

The purpose of this stage was to introduce the model to:

  • —Liujgoj Cantonese Orthography
  • —Cantonese phonological structure
  • —Cantonese vocabulary and expressions
  • —Liujgoj spelling patterns
  • —Romanized Cantonese text
  • —Cantonese sentence structure

An important design decision was that Cantonese syllables were not pre-registered as additional tokenizer tokens before training.

The model therefore learned Liujgoj using the original tokenizer and vocabulary of Qwen2.5-Omni.

The Cantonese syllables were not pre-registered as tokens before training.

This simplified the training pipeline and avoided introducing thousands of newly initialized token embeddings.

Despite the relatively limited CPT dataset, the model successfully developed the ability to understand and generate structured Liujgoj Cantonese.


Stage 2 — Supervised Fine-Tuning (SFT)

The CPT model was then fine-tuned using instruction-based training data.

This stage further developed the model's ability to perform structured Liujgoj-related tasks, including:

  • —Cantonese Chinese characters → Liujgoj
  • —Liujgoj → Cantonese Chinese characters
  • —Liujgoj syllable segmentation
  • —Liujgoj structural analysis
  • —Cantonese semantic understanding

The purpose of the SFT stage was not simply to teach the model Liujgoj, but to organize the linguistic knowledge acquired during CPT into usable instruction-following capabilities.

Conceptually:

text
CPT
↓
Learn the Liujgoj Cantonese language space

SFT
↓
Learn how to use that knowledge for structured tasks

ASR
↓
Learn to map spoken Cantonese into Liujgoj Cantonese

Stage 3 — Automatic Speech Recognition

The SFT model was subsequently fine-tuned using paired Cantonese audio and Liujgoj transcription data.

The objective is:

text
Audio
  ↓
Qwen2.5-Omni Audio Understanding
  ↓
Liujgoj Cantonese Language Space
  ↓
Liujgoj Cantonese Transcription

Because the model had already been exposed to Liujgoj through CPT and SFT before ASR training, the ASR stage focuses primarily on learning the relationship between:

Cantonese speech and Liujgoj text

rather than learning the entire Liujgoj orthographic system from scratch.


📊 Training Data

The ASR model was fine-tuned using paired Cantonese speech and Liujgoj Romanized Cantonese transcription data.

  • —Number of Training Examples: 35,032 audio–text pairs
  • —Total Audio Duration: Approximately 18 hours of audio
  • —Audio Format: 16 kHz, mono WAV
  • —Transcription System: Liujgoj Cantonese Orthography (溜歌粵語羅馬字)
  • —Data Source: Cantonese movies

Each training example consists of:

text
Cantonese Audio
        ↓
Liujgoj Cantonese Transcription

The training objective is therefore to directly map spoken Cantonese audio to structured Liujgoj Romanized Cantonese.

text
Audio
  ↓
ASR Model
  ↓
Liujgoj Cantonese

The training data contains natural Cantonese speech extracted from Cantonese movie material, providing conversational speech, diverse linguistic contexts, and naturally occurring Cantonese expressions.

Unlike conventional Cantonese ASR datasets that typically use Chinese characters as transcription targets, the text labels in this dataset are represented directly in Liujgoj Cantonese Orthography.

This allows the model to learn the following mapping directly:

text
Spoken Cantonese
        ↓
Liujgoj syllables
        +
Tone letters
        +
Word boundaries
        +
Sentence structure

Training Data Summary

ItemDetails
Training examples35,032
Total audio duration~18 hours
Audio sampling rate16 kHz
Audio channelsMono
Audio formatWAV
Speech languageCantonese
Transcription targetLiujgoj Cantonese
Data sourceCantonese movies

Key Features

🎙️ Direct Cantonese Speech Recognition

Transcribes spoken Cantonese audio directly into Liujgoj Cantonese.

text
Cantonese Speech

        ↓

Liujgoj Cantonese

🔤 Romanized Cantonese Output

The model does not require Chinese characters as an intermediate transcription target.

Instead, it directly generates Romanized Cantonese according to the Liujgoj orthographic system.

This makes the model suitable for:

  • —Cantonese speech corpora
  • —Romanized Cantonese datasets
  • —Cantonese language research
  • —Phonological analysis
  • —Cantonese language learning
  • —Speech data annotation
  • —Cantonese TTS pipelines
  • —Large-scale Cantonese audio transcription

🔊 Tone-Aware Output

Liujgoj represents Cantonese tones using letters.

The ASR model is therefore trained not only to recognize Cantonese syllables but also to generate the corresponding tone information.

For example, the transcription output contains structured forms such as:

text
Cinxminh ngoq saamj gaaj coux cej aax

where the Romanized output includes syllable structure and tone-related orthographic information.


✂️ Structural Word Segmentation

Unlike character-level transcription systems, Liujgoj uses spaces to represent word boundaries.

The model therefore learns to generate:

text
word word word

instead of producing an unsegmented sequence of Romanized syllables.

This enables the output to preserve both:

  • —phonological information
  • —lexical structure

🧠 Context-Aware Transcription

The model inherits language understanding capabilities from the Qwen2.5-Omni architecture and is further adapted through Liujgoj CPT and SFT.

This allows the model to use linguistic context when generating Romanized Cantonese.

The intended pipeline is therefore not simply:

text
Audio → Phonetic Symbols

but:

text
Audio
  ↓
Acoustic Understanding
  +
Liujgoj Cantonese Language Knowledge
  ↓
Structured Liujgoj Transcription

Example

Audio File

測試音訊選自《鹿鼎記》開場白(來源:CanCLID/zoengjyutgaai):

text
001_001.opus, 001_002.opus, 001_003.opus

Model Output

text
Waahzvj hair Cingjciux ge Hongjheij cojninr, dungjtinj ge mouxhoengjyanx, bakj fungj yvx douj, munq deih bingjsoengj.
Hair Gongjnaamx kaauganr hoirbinh yatj tiux daaihlouh zijsoengh, yauq yatj deoi cingjbingj, saurzath doujcoengj, ngath zvh catj gaaj caurcej, hoeng zvh bakjfongj haangxganr.
Cinxminh ngoq saamj gaaj coux cej aax, muiq gaaj zauh wanr zvh yatj go naamxzir, douj haih svjsangj daarbaanhnaaix.

Reference Cantonese Text

text
話説喺清朝嘅康熙初年,冬天嘅某一日,北風如刀,滿地冰霜啊。

喺江南靠近海邊一條大路之上,有一隊清兵,手執刀槍,押住七架囚車,向住北方行緊。

前便嗰三架囚車啊,每架就韞住一個男子,都係書生打扮。

This example demonstrates the model's ability to generate structured Liujgoj Cantonese directly from spoken Cantonese audio.

As the model is still under development, transcription errors may occur, particularly in difficult acoustic conditions, rare vocabulary, proper nouns, and long-form speech.


Model Lineage

text
Qwen2.5-Omni-7B
        │
        ▼
Liujgoj-Cantonese-Qwen2.5-Omni-7B-CPT
        │
        │  Continued Pretraining
        │  ~14M training data
        │  Original tokenizer
        │  No Cantonese syllable tokens added
        ▼
Liujgoj-Cantonese-Qwen2.5-Omni-7B-SFT
        │
        │  Instruction Fine-Tuning
        │  Cantonese ↔ Liujgoj
        │  Syllable segmentation
        │  Structural understanding
        ▼
Liujgoj-Cantonese-Qwen2.5-Omni-7B-ASR
        │
        │  Audio + Liujgoj transcription
        ▼
Cantonese Speech → Liujgoj Cantonese

Why This Training Pipeline?

The central idea behind this project is to separate the learning process into three stages.

1. Learn the language

text
CPT

The model learns the Liujgoj Cantonese orthographic system and Cantonese linguistic patterns.

2. Learn the tasks

text
SFT

The model learns how to perform structured Liujgoj-related tasks through instructions.

3. Learn the speech mapping

text
ASR

The model learns to connect spoken Cantonese audio with the Liujgoj language space already established during the previous stages.

This results in the overall architecture:

text
Language Knowledge
        +
Instruction Following
        +
Audio Understanding
        ↓
Liujgoj Cantonese ASR

Quickstart

Installation

Install the required dependencies:

bash
pip install torch transformers librosa accelerate

Load the Model

python
import torch
from transformers import AutoProcessor, AutoModel

model_id = "Yvthyvq/Liujgoj-Cantonese-Qwen2.5-Omni-7B-ASR"

processor = AutoProcessor.from_pretrained(
    model_id,
    trust_remote_code=True
)

model = AutoModel.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True
)

model.eval()

Prepare Audio

python
import librosa

audio, sr = librosa.load(
    "your_audio.wav",
    sr=16000
)

Run Inference

The exact inference prompt and audio input format follow the Qwen2.5-Omni multimodal processing pipeline.

A typical instruction is:

text
Please transcribe the following Cantonese audio into Liujgoj Cantonese.

Intended Use

This model is intended for research and development involving:

  • —Cantonese Automatic Speech Recognition
  • —Romanized Cantonese transcription
  • —Liujgoj Cantonese Orthography
  • —Cantonese speech corpus construction
  • —Cantonese phonological research
  • —Cantonese language learning
  • —Speech-to-text pipelines
  • —Cantonese TTS data preparation

Current Status

This model represents an experimental stage of the Liujgoj Cantonese multimodal model family.

The training pipeline has demonstrated that Liujgoj Cantonese capabilities can be developed through:

text
Original Qwen Tokenizer
        ↓
Liujgoj CPT
        ↓
Liujgoj SFT
        ↓
Liujgoj ASR

without pre-registering Cantonese syllables as additional tokenizer tokens.

Future improvements may include:

  • —Larger Cantonese speech datasets
  • —More diverse speakers
  • —Additional acoustic domains
  • —Long-form Cantonese transcription
  • —Improved tone accuracy
  • —Improved word segmentation
  • —Automatic punctuation
  • —Comparison with Whisper-based Cantonese ASR
  • —Comparison with SenseVoice-based Cantonese ASR
  • —Further integration with Liujgoj Cantonese TTS systems

Related Models