CoolFace
Modelpublic

tintitu/speech_paraformer-large-contextual_asr_nat-zh-cn-16k-common-vocab8404-onnx

sourceHugging Faceupdated 17d agoView on Hugging Face
0likes
Model Card

Paraformer-Large-Contextual 热词增强语音识别模型 (ONNX 量化版)

面向中文自动语音识别(ASR)的端到端非自回归大模型(集成上下文热词偏置网络 Contextual Biasing Network),当前目录包含 ONNX 量化主模型、热词偏置嵌入网络、分词词典、分词词表及解码配置。

模型信息

  • —模型名称:Paraformer-Large-Contextual ASR (NAT-ZH-CN-16k-Common-Vocab8404)
  • —适用任务:离线语音识别(ASR,支持动态热词/专有名词增强)
  • —模型架构:非自回归 Transformer + Contextual Biasing Network
  • —支持语言:中文普通话(zh-cn)
  • —词表规模:8404 常用汉字与特殊 Token
  • —特色能力:支持在推理时动态注入热词列表(如人名、地名、行业术语、自定义品牌词),显著提升特定词汇的召回率与准确率。
  • —主要文件:
  • —model_quant.onnx:经过 INT8 静态量化的声学主干推理模型
  • —model_eb.onnx:热词嵌入偏置网络(Embedding Biaser)
  • —seg_dict:热词切词与分词辅助字典
  • —tokens.json:字符与词元 Token 映射词表
  • —config.yaml:模型推理与热词协同解码配置
  • —am.mvn:声学特征 CMVN(倒谱均值方差归一化)统计文件

硬件资源门槛与性能预期 (CPU)

  • —内存常驻与峰值:约 1.1 GB ~ 1.6 GB RAM(包含热词偏置网络与分词字典常驻,推理峰值约 1.4 GB)
  • —性能预期 (RTF):典型 8 核 CPU 上,无热词或百词以内小热词表时 RTF 预期约为 0.12 ~ 0.22(处理 10 秒音频耗时约 1.2 ~ 2.2 秒)
  • —推荐配置:4 核及以上 CPU、8 GB 及以上系统物理内存

音频输入规范

  • —采样率要求:16000 Hz(16 kHz)
  • —声道配置:单声道(Mono)
  • —音频格式:16-bit PCM 或归一化浮点格式(数值范围 [-1.0, 1.0])
  • —特征提取:默认提取 80 维 Fbank 特征,配合 CMVN 归一化输入

推荐流水线协作

  • —前置模块:推荐搭配 speech_fsmn_vad_zh-cn-16k-common-onnx 进行语音活动检测与长音频断句,滤除环境静音并提高识别效率。
  • —后置模块:识别输出为纯文字流,推荐后置串联标点恢复模型(如 punc_ct-transformer_zh-cn-common-vocab272727-onnx)完成自动标点补充。
  • —热词调用:热词输入通常以空格或换行分隔;若不传入热词,模型表现与标准版 Paraformer-Large 基本一致。

上游来源

本目录包含的是上游模型导出并量化后的 ONNX 推理版本,版权归原作者及阿里巴巴达摩院/通义实验室所有。

许可证与再分发

遵循上游 FunASR 与 ModelScope 开源协议(Apache-2.0 / ModelScope 社区许可)。商业集成或二次分发时,请保留上游版权声明、许可说明及原始出处链接。

文件校验清单 (SHA-256)

下载后请核对文件大小和 SHA256 校验和。

文件名大小 (字节)SHA-256 校验和
am.mvn11,20329b3c740a2c0cfc6b308126d31d7f265fa2be74f3bb095cd2f143ea970896ae5
config.yaml2,5321d9057edeaba9e131cb98f26011606497cf3af187d8943525ddb5ee36c836b1b
model_eb.onnx25,618,359d31446a5af664291a2922cca253a4200a523f347d6fc3cb1bff356bf60a116b6
model_quant.onnx871,251,660f404e6eb532b54fd95761e2b4be4ed1998e8cff3cb3b930a9bee1f2d556e5035
seg_dict8,287,83459a2ef803a3f1648ad03a2e1480db1c1ee0c0d7dc4ef4dbd16cea33944329022
tokens.json93,6762b20c2b12572d682afff84ce1c8d560f67b8b32a4c1f21567411d141ed352127

Paraformer-Large-Contextual ASR Model (ONNX Quantized) (English Documentation)

An end-to-end non-autoregressive speech recognition model with contextual hotword biasing for Mandarin Chinese ASR. This directory contains quantized ONNX weights, embedding biaser network, segmentation dictionary, vocabulary mappings, and configuration files.

Model Information

  • —Model Name: Paraformer-Large-Contextual ASR (NAT-ZH-CN-16k-Common-Vocab8404)
  • —Task: Offline Automatic Speech Recognition (ASR) with dynamic hotword enhancement
  • —Architecture: Non-Autoregressive Transformer + Contextual Biasing Network
  • —Language: Mandarin Chinese (zh-cn)
  • —Vocabulary Size: 8,404 Chinese characters and special tokens
  • —Key Feature: Supports runtime hotword injection (e.g., entity names, specialized terms, branded keywords) to substantially improve keyword recall and precision.
  • —Primary Files:
  • —model_quant.onnx: INT8 statically quantized acoustic inference backbone
  • —model_eb.onnx: Embedding biaser network for contextual hotword injection
  • —seg_dict: Word segmentation dictionary for hotword tokenization
  • —tokens.json: Character and token mapping dictionary
  • —config.yaml: Model configuration and biaser decoding parameters
  • —am.mvn: Cepstral mean and variance normalization (CMVN) statistics

Hardware Footprint & Sizing (CPU)

  • —Memory Footprint (RAM): ~1.1 GB to 1.6 GB (Including biaser model and dictionary cache, peak ~1.4 GB)
  • —Expected RTF (Real-Time Factor): RTF of ~0.12 - 0.22 on typical 8-core CPUs with standard hotword lists (processing a 10s audio clip takes ~1.2s - 2.2s)
  • —Recommended System: 4+ CPU cores, 8+ GB RAM

Audio Input Specifications

  • —Sampling Rate: 16000 Hz (16 kHz)
  • —Channels: Single channel (Mono)
  • —Format: 16-bit PCM or normalized float array within [-1.0, 1.0]
  • —Feature Extraction: 80-dimensional log Mel-filterbank features with CMVN

Recommended Pipeline Integration

  • —Upstream Preprocessing: Recommended pairing with speech_fsmn_vad_zh-cn-16k-common-onnx for voice activity detection and sentence segmentation.
  • —Downstream Postprocessing: Paraformer outputs raw text without punctuation. It is recommended to cascade a punctuation restoration model (such as punc_ct-transformer_zh-cn-common-vocab272727-onnx).
  • —Hotword Usage: Hotword inputs are typically supplied as whitespace- or newline-separated lists. When no hotwords are provided, performance is equivalent to standard Paraformer-Large.

Upstream Sources

This directory provides a quantized ONNX inference distribution derived from upstream checkpoints. All copyrights belong to the original authors and Alibaba DAMO Academy / Tongyi Lab.

Licensing and Redistribution

Governed by the upstream FunASR and ModelScope licensing terms (Apache-2.0 / ModelScope Community License). Please retain upstream copyright statements, notices, and source links in downstream redistributions.

File Verification Manifest (SHA-256)

Verify file sizes and SHA-256 checksums after downloading.

FilenameSize (Bytes)SHA-256 Checksum
am.mvn11,20329b3c740a2c0cfc6b308126d31d7f265fa2be74f3bb095cd2f143ea970896ae5
config.yaml2,5321d9057edeaba9e131cb98f26011606497cf3af187d8943525ddb5ee36c836b1b
model_eb.onnx25,618,359d31446a5af664291a2922cca253a4200a523f347d6fc3cb1bff356bf60a116b6
model_quant.onnx871,251,660f404e6eb532b54fd95761e2b4be4ed1998e8cff3cb3b930a9bee1f2d556e5035
seg_dict8,287,83459a2ef803a3f1648ad03a2e1480db1c1ee0c0d7dc4ef4dbd16cea33944329022
tokens.json93,6762b20c2b12572d682afff84ce1c8d560f67b8b32a4c1f21567411d141ed352127