CoolFace
Modelpublic

ILRDF/whisper-large-v2-formosan-lang-tokens

sourceHugging Facecc-by-nc-4.0updated 19d agoView on Hugging Face
0likes10downloads
Model Card

Model Card for whisper-large-v2-formosan-lang-tokens

This model is a fine-tuned version of ILRDF/whisper-large-v2-formosan-all (itself based on openai/whisper-large-v2) for automatic speech recognition across 42 Formosan indigenous language dialects of Taiwan.

Unlike the base model - which conditions every dialect on a single placeholder Whisper language token (id, Indonesian) - this model adds one real Whisper-style language token per dialect (e.g. <|ami-x-frng|>), the same mechanism Whisper itself uses for its original ~99 languages, just extended to 42 dialects it was never pretrained on. Each dialect's own added token replaces the usual <|lang|> slot in the decoder prefix, so the model is told exactly which dialect it's transcribing instead of having to infer it (or default to one shared representation for all of them).

Supported dialects

42 dialects across 15 language families of Taiwan's Formosan indigenous languages. Each has its own added language token, replacing the <|lang|> slot in the decoder prefix <|startoftranscript|><|lang|><|transcribe|><|notimestamps|>.

<details> <summary>Full list of 42 dialects and their tokens (click to expand)</summary>

Tokenlang_codeFamily (EN)Family (中文)Dialect (中文)
`<\ami-x-frng\>` (id 51865)ami-x-frngAmis阿美語馬蘭阿美語
`<\ami-x-iams\>` (id 51866)ami-x-iamsAmis阿美語南勢阿美語
`<\ami-x-pld\>` (id 51867)ami-x-pldAmis阿美語恆春阿美語
`<\ami-x-pswl\>` (id 51868)ami-x-pswlAmis阿美語海岸阿美語
`<\ami-x-skl\>` (id 51869)ami-x-sklAmis阿美語秀姑巒阿美語
`<\bnn-x-bkh\>` (id 51870)bnn-x-bkhBunun布農語卡群布農語
`<\bnn-x-bnz\>` (id 51871)bnn-x-bnzBunun布農語巒群布農語
`<\bnn-x-isbk\>` (id 51872)bnn-x-isbkBunun布農語郡群布農語
`<\bnn-x-td\>` (id 51873)bnn-x-tdBunun布農語卓群布農語
`<\bnn-x-vtn\>` (id 51874)bnn-x-vtnBunun布農語丹群布農語
`<\ckv\>` (id 51875)ckvKebalan噶瑪蘭語噶瑪蘭語
`<\dru-x-kgdv\>` (id 51876)dru-x-kgdvRukai魯凱語多納魯凱語
`<\dru-x-lbw\>` (id 51877)dru-x-lbwRukai魯凱語大武魯凱語
`<\dru-x-ngdr\>` (id 51878)dru-x-ngdrRukai魯凱語霧臺魯凱語
`<\dru-x-opnh\>` (id 51879)dru-x-opnhRukai魯凱語萬山魯凱語
`<\dru-x-tldr\>` (id 51880)dru-x-tldrRukai魯凱語茂林魯凱語
`<\dru-x-trmk\>` (id 51881)dru-x-trmkRukai魯凱語東魯凱語
`<\pwn-x-kcdsn\>` (id 51882)pwn-x-kcdsnPaiwan排灣語東排灣語
`<\pwn-x-pnvn\>` (id 51883)pwn-x-pnvnPaiwan排灣語中排灣語
`<\pwn-x-vnrn\>` (id 51884)pwn-x-vnrnPaiwan排灣語北排灣語
`<\pwn-x-ynvl\>` (id 51885)pwn-x-ynvlPaiwan排灣語南排灣語
`<\pyu-x-ksvk\>` (id 51886)pyu-x-ksvkPinuyumayan卑南語建和卑南語
`<\pyu-x-ktrp\>` (id 51887)pyu-x-ktrpPinuyumayan卑南語知本卑南語
`<\pyu-x-mkzy\>` (id 51888)pyu-x-mkzyPinuyumayan卑南語西群卑南語
`<\pyu-x-pym\>` (id 51889)pyu-x-pymPinuyumayan卑南語南王卑南語
`<\ssf\>` (id 51890)ssfThau邵語邵語
`<\sxr\>` (id 51891)sxrHla'alua拉阿魯哇語拉阿魯哇語
`<\szy\>` (id 51892)szySakizaya撒奇萊雅語撒奇萊雅語
`<\tao\>` (id 51893)taoYami雅美語雅美語
`<\tay-x-cql\>` (id 51894)tay-x-cqlTayal泰雅語四季泰雅語
`<\tay-x-kls\>` (id 51895)tay-x-klsTayal泰雅語宜蘭澤敖利泰雅語
`<\tay-x-mtuw\>` (id 51896)tay-x-mtuwTayal泰雅語汶水泰雅語
`<\tay-x-plngw\>` (id 51897)tay-x-plngwTayal泰雅語萬大泰雅語
`<\tay-x-sql\>` (id 51898)tay-x-sqlTayal泰雅語賽考利克泰雅語
`<\tay-x-sul\>` (id 51899)tay-x-sulTayal泰雅語澤敖利泰雅語
`<\trv-x-td\>` (id 51900)trv-x-tdTruku/Seediq太魯閣語/賽德克語都達賽德克語
`<\trv-x-tgdy\>` (id 51901)trv-x-tgdyTruku/Seediq太魯閣語/賽德克語德固達雅賽德克語
`<\trv-x-trk\>` (id 51902)trv-x-trkTruku/Seediq太魯閣語/賽德克語德鹿谷賽德克語
`<\trv-x-truku\>` (id 51903)trv-x-trukuTruku/Seediq太魯閣語/賽德克語太魯閣語
`<\tsu\>` (id 51904)tsuCou鄒語鄒語
`<\xnb\>` (id 51905)xnbKanakanavu卡那卡那富語卡那卡那富語
`<\xsy\>` (id 51906)xsySaySiyat賽夏語賽夏語

</details>

Results

Compared against the base model, `ILRDF/whisper-large-v2-formosan-all`, on the eval split of every config of formospeech/klokah and formospeech/ithuan_formosan (45 dialect/dataset pairs total). Numbers below are normalized WER/CER (lowercased, light punctuation stripped - see "Text normalization" in the details below for exactly what that means).

Macro-average (mean of each dialect's own metric) across all 45, and separately per dataset:

ScopeModelNormalized WERNormalized CER
All 45 (pooled)baseline10.312.82
All 45 (pooled)ours7.472.11
formospeech/klokah only (n=42)baseline10.482.91
formospeech/klokah only (n=42)ours7.672.19
formospeech/ithuan_formosan only (n=3)baseline8.041.59
formospeech/ithuan_formosan only (n=3)ours4.751.01

This model beats the baseline's normalized WER/CER on 44/45 and 43/45 dialects respectively (the handful of losses are within ~0.1-0.5 points); on raw, non-normalized WER/CER it wins 45/45.

Noise robustness (MUSAN noise mixed in at the target SNR before transcription, same 45 pairs, macro-averaged):

ConditionNormalized WERNormalized CER
Clean7.472.11
15 dB SNR11.483.17
10 dB SNR14.754.30

<details> <summary>Full per-dialect results & methodology (click to expand)</summary>

Every number above and in the per-dialect table below is broken down per dataset, per config (dialect), and per split; the pooled figures are macro-averages computed from those per-dialect numbers, not a single blended metric. The model-index metadata in this README's YAML frontmatter carries the same breakdown machine-readably: one results entry per (dataset, config, split) triple, 45 in total, each with its own dataset.type/dataset.name/dataset.config/dataset.split and wer/cer (raw and normalized) metrics - this is what powers the auto-rendered results box on this model's Hub page. Both datasets only ship train/eval splits, and eval is the one used everywhere here - there is no separate held-out test split.

Note on the model-index format used here vs. Hugging Face's Evaluation Results docs: that page documents a newer, separate mechanism (.eval_results/*.yaml files, tied to a dataset repo that's registered as a Hub Benchmark with its own eval.yaml) built for community-submitted, semi-verified leaderboard-style results. Neither formospeech/klokah nor formospeech/ithuan_formosan is registered as a Benchmark, so that mechanism doesn't apply to this model card. The model-index: block above is the older, more widely-supported convention (the same one used to render the standard "Results" table on model pages across the Hub, including the sibling formospeech/whisper-large-v2-taiwanese-hakka-v1 model card) and is what's actually used here.

Text normalization

Both models' predictions and references are lowercased, stripped of ., ,, !, ?, ; and have runs of whitespace collapsed to a single space before computing the normalized metrics - deliberately a small, fixed whitelist rather than stripping every Unicode punctuation/symbol character (e.g. transformers' BasicTextNormalizer), because two characters that look like punctuation are actually meaningful orthography in this data and would otherwise get silently corrupted:

  • —`:` (colon) marks vowel length in some dialects (Amis, Saisiyat) - e.g. bae:iw, sapi:ihin. A corpus-wide scan found it directly between two letters (i.e. part of a word, not separating sentences) 77-95% of the time in the dialects that use it.
  • —U+2303 (`⌃`) is a meaningful orthographic marker that also occurs in the data.

Neither is stripped. Every other punctuation character actually appearing in the training corpora was confirmed (via the same scan) to never occur mid-word, so stripping just those five is safe.

Noise robustness methodology

MUSAN noise-subset clips (bilguun/musan-noise, 930 clips) are additively mixed in at a target SNR before transcription, same 45 dialect/dataset pairs as above. Raw (non-normalized) WER/CER for the same three conditions:

ConditionWERCER
Clean10.382.64
15 dB SNR15.053.85
10 dB SNR18.635.08

All four metrics (WER/CER, raw and normalized) degrade monotonically as SNR decreases, as expected. The Rukai (dru-*) dialects are consistently the most noise-sensitive of the 42, with dru-x-opnh and dru-x-kgdv losing 15-20 points of absolute WER at 10dB SNR - worth keeping in mind for any downstream use in noisy conditions. This is one condition among many possible robustness probes (there's no single agreed-upon standard for which MUSAN subset(s) or SNR range to use for ASR robustness evaluation specifically - MUSAN itself was originally designed for training-time augmentation recipes, not a fixed eval protocol). Results here should be read as "robustness to point-source background noise at moderate SNR," not a full robustness certification.

Per-dialect breakdown
lang_codeFamilyDialectDatasetnorm_WERBaseline norm_WERnorm_CERBaseline norm_CER15dB norm_WER10dB norm_WER
ami-x-frngAmis馬蘭阿美語klokah4.635.681.61.897.8210.02
ami-x-iamsAmis南勢阿美語klokah6.0915.261.3210.298.711.82
ami-x-pldAmis恆春阿美語klokah5.496.91.071.339.1813.38
ami-x-pswlAmis海岸阿美語klokah3.313.780.610.736.08.79
ami-x-sklAmis秀姑巒阿美語ithuan_formosan5.039.551.081.637.049.55
ami-x-sklAmis秀姑巒阿美語klokah4.656.091.061.367.799.68
bnn-x-bkhBunun卡群布農語klokah16.7119.665.185.2922.4526.75
bnn-x-bnzBunun巒群布農語klokah7.1921.752.033.3412.1316.19
bnn-x-isbkBunun郡群布農語klokah7.669.641.681.9612.7416.34
bnn-x-tdBunun卓群布農語klokah10.7712.82.312.4814.3116.75
bnn-x-vtnBunun丹群布農語klokah11.6917.713.84.3216.2619.6
ckvKebalan噶瑪蘭語klokah4.585.091.221.377.619.64
dru-x-kgdvRukai多納魯凱語klokah13.6315.053.413.7721.7728.0
dru-x-lbwRukai大武魯凱語klokah9.2110.522.853.115.1220.03
dru-x-ngdrRukai霧臺魯凱語klokah8.510.663.162.7114.9720.0
dru-x-opnhRukai萬山魯凱語klokah12.6815.93.113.4525.6632.16
dru-x-tldrRukai茂林魯凱語klokah34.8854.018.6414.8242.3446.6
dru-x-trmkRukai東魯凱語klokah11.6113.062.572.9417.8923.04
pwn-x-kcdsnPaiwan東排灣語klokah8.710.11.893.4412.5314.62
pwn-x-pnvnPaiwan中排灣語klokah8.19.351.82.1311.2913.84
pwn-x-vnrnPaiwan北排灣語klokah8.69.72.132.3413.3216.97
pwn-x-ynvlPaiwan南排灣語klokah11.3212.562.713.0315.1518.24
pyu-x-ksvkPinuyumayan建和卑南語klokah2.983.420.650.836.5510.11
pyu-x-ktrpPinuyumayan知本卑南語klokah3.493.970.790.856.9210.24
pyu-x-mkzyPinuyumayan西群卑南語klokah3.854.521.291.458.0411.26
pyu-x-pymPinuyumayan南王卑南語klokah2.248.070.611.855.097.7
ssfThau邵語klokah4.725.311.661.847.249.62
sxrHla'alua拉阿魯哇語klokah10.3410.881.631.7516.5320.41
szySakizaya撒奇萊雅語klokah4.185.611.211.587.010.26
taoYami雅美語klokah9.6110.583.03.2914.318.18
tay-x-cqlTayal四季泰雅語klokah5.056.231.712.018.6311.76
tay-x-klsTayal宜蘭澤敖利泰雅語klokah3.034.431.291.656.29.13
tay-x-mtuwTayal汶水泰雅語klokah7.859.162.993.513.917.93
tay-x-plngwTayal萬大泰雅語klokah6.358.151.521.9810.7813.88
tay-x-sqlTayal賽考利克泰雅語klokah4.155.551.321.827.6411.67
tay-x-sulTayal澤敖利泰雅語klokah3.514.41.241.547.4511.38
trv-x-tdTruku/Seediq都達賽德克語klokah3.874.131.461.476.269.75
trv-x-tgdyTruku/Seediq德固達雅賽德克語ithuan_formosan6.639.941.532.197.738.29
trv-x-tgdyTruku/Seediq德固達雅賽德克語klokah4.814.982.322.458.110.72
trv-x-trkTruku/Seediq德鹿谷賽德克語klokah6.836.624.284.298.039.96
trv-x-trukuTruku/Seediq太魯閣語ithuan_formosan2.584.640.420.942.063.61
trv-x-trukuTruku/Seediq太魯閣語klokah3.83.891.861.85.537.32
tsuCou鄒語klokah8.949.863.143.3112.5915.61
xnbKanakanavu卡那卡那富語klokah3.665.811.141.725.517.8
xsySaySiyat賽夏語klokah8.7829.182.855.1212.3215.02

</details>

Usage

Access and Authentication

This model is hosted as a gated Hugging Face repository. Before using it:

  1. 1.Visit the model page and request access.
  2. 2.Log in with the same Hugging Face account that has been granted access.
  3. 3.Authenticate your local environment with a Hugging Face access token.

A read token is sufficient for inference.

bash
pip install -U huggingface_hub
hf auth login

Alternatively, you can provide the token through the HF_TOKEN environment variable:

bash
export HF_TOKEN=hf_xxx

Do not hard-code your Hugging Face token in scripts, notebooks, or public repositories.

If you see an error such as Cannot access gated repo, make sure that:

  • —your Hugging Face account has been granted access to this model;
  • —hf auth whoami shows the expected account;
  • —HF_HUB_DISABLE_IMPLICIT_TOKEN is not set.

Run model

Because each dialect uses its own added language token instead of one of Whisper's built-in languages, the usual pipeline(..., generate_kwargs={"language": ...}) shortcut does not apply here - the convenience language=/task= arguments only recognize Whisper's original ~99 languages. Instead, force the dialect's token directly via decoder_input_ids:

python
import json
import torch
from huggingface_hub import hf_hub_download
from transformers import WhisperForConditionalGeneration, WhisperProcessor

model_id = "formospeech/whisper-large-v2-formosan-lang-tokens"
device = "cuda:0" if torch.cuda.is_available() else "cpu"
torch_dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32

processor = WhisperProcessor.from_pretrained(model_id, language="id", task="transcribe")
model = WhisperForConditionalGeneration.from_pretrained(model_id, torch_dtype=torch_dtype).to(device)

# lang_code -> added token id, e.g. "ami-x-frng" -> 51865 (see "Supported dialects" above)
lang_code_to_token_id = json.load(open(hf_hub_download(model_id, "lang_code_to_token_id.json")))

def transcribe(audio_array, sampling_rate, lang_code):
    decoder_start_id, _, transcribe_id, notimestamps_id = processor.tokenizer.prefix_tokens
    dialect_token_id = lang_code_to_token_id[lang_code]
    # Passing decoder_input_ids to Whisper's generate() is taken verbatim (it does NOT
    # auto-prepend decoder_start_token_id the way the base GenerationMixin does), so all
    # four prefix tokens must be given explicitly here.
    decoder_input_ids = torch.tensor(
        [[decoder_start_id, dialect_token_id, transcribe_id, notimestamps_id]], device=device
    )
    input_features = processor.feature_extractor(
        audio_array, sampling_rate=sampling_rate, return_tensors="pt"
    ).input_features.to(device, dtype=torch_dtype)
    with torch.no_grad():
        generated_ids = model.generate(input_features, decoder_input_ids=decoder_input_ids, max_length=225)
    return processor.tokenizer.decode(generated_ids[0], skip_special_tokens=True)

# example: audio_array is a float32 numpy array at 16kHz
# transcribe(audio_array, 16000, "ami-x-frng")

Training process

The training of the model was performed with the following hyperparameters:

  • —Hardware: 4x NVIDIA RTX A5000
  • —Per-device batch size: 2
  • —Gradient accumulation steps: 64
  • —Effective batch size: 512
  • —Total training steps: 1159 (1 epoch)
  • —Learning rate: 1e-4
  • —Warmup ratio: 0.1
  • —Precision: bf16
  • —Optimizer: adamw_torch
  • —LR scheduler type: linear
  • —Initialization: base model's weights (ILRDF/whisper-large-v2-formosan-all), with the 42 new dialect token embeddings each initialized from that model's own id (Indonesian) token embedding - a linguistically related Austronesian language, per Whisper maintainer guidance on adding new languages

The learning rate was chosen empirically via a short LR probe (3e-5 / 1e-4 / 3e-4 candidates, ~100 steps each) before committing to a full run - 1e-4 gave the best eval WER among the three.

Training data

Every config (dialect) of every dataset's train split was loaded and concatenated:

  • —formospeech/ntu_formosan_corpus
  • —formospeech/ilrdf_dicts
  • —formospeech/klokah
  • —formospeech/ithuan_formosan

Notes

  • —This release contains inference files only. Optimizer states and trainer checkpoints are intentionally excluded.
  • —lang_code_to_token_id.json (included in this repo) is required to look up the correct token id per dialect - see the usage example above.