tintitu/speech_paraformer-large-contextual_asr_nat-zh-cn-16k-common-vocab8404-onnx
Paraformer-Large-Contextual 热词增强语音识别模型 (ONNX 量化版)
面向中文自动语音识别(ASR)的端到端非自回归大模型(集成上下文热词偏置网络 Contextual Biasing Network),当前目录包含 ONNX 量化主模型、热词偏置嵌入网络、分词词典、分词词表及解码配置。
模型信息
- 模型名称:Paraformer-Large-Contextual ASR (NAT-ZH-CN-16k-Common-Vocab8404)
- 适用任务:离线语音识别(ASR,支持动态热词/专有名词增强)
- 模型架构:非自回归 Transformer + Contextual Biasing Network
- 支持语言:中文普通话(zh-cn)
- 词表规模:8404 常用汉字与特殊 Token
- 特色能力:支持在推理时动态注入热词列表(如人名、地名、行业术语、自定义品牌词),显著提升特定词汇的召回率与准确率。
- 主要文件:
model_quant.onnx:经过 INT8 静态量化的声学主干推理模型model_eb.onnx:热词嵌入偏置网络(Embedding Biaser)seg_dict:热词切词与分词辅助字典tokens.json:字符与词元 Token 映射词表config.yaml:模型推理与热词协同解码配置am.mvn:声学特征 CMVN(倒谱均值方差归一化)统计文件
硬件资源门槛与性能预期 (CPU)
- 内存常驻与峰值:约 1.1 GB ~ 1.6 GB RAM(包含热词偏置网络与分词字典常驻,推理峰值约 1.4 GB)
- 性能预期 (RTF):典型 8 核 CPU 上,无热词或百词以内小热词表时 RTF 预期约为 0.12 ~ 0.22(处理 10 秒音频耗时约 1.2 ~ 2.2 秒)
- 推荐配置:4 核及以上 CPU、8 GB 及以上系统物理内存
音频输入规范
- 采样率要求:16000 Hz(16 kHz)
- 声道配置:单声道(Mono)
- 音频格式:16-bit PCM 或归一化浮点格式(数值范围 [-1.0, 1.0])
- 特征提取:默认提取 80 维 Fbank 特征,配合 CMVN 归一化输入
推荐流水线协作
- 前置模块:推荐搭配
speech_fsmn_vad_zh-cn-16k-common-onnx进行语音活动检测与长音频断句,滤除环境静音并提高识别效率。 - 后置模块:识别输出为纯文字流,推荐后置串联标点恢复模型(如
punc_ct-transformer_zh-cn-common-vocab272727-onnx)完成自动标点补充。 - 热词调用:热词输入通常以空格或换行分隔;若不传入热词,模型表现与标准版 Paraformer-Large 基本一致。
上游来源
- 上游开源项目:`alibaba-damo-academy/FunASR`
- ModelScope 官方仓库:`damo/speech_paraformer-large-contextual_asr_nat-zh-cn-16k-common-vocab8404-onnx`
- 推理引擎兼容:ONNX Runtime、sherpa-onnx、FunASR Runtime
本目录包含的是上游模型导出并量化后的 ONNX 推理版本,版权归原作者及阿里巴巴达摩院/通义实验室所有。
许可证与再分发
遵循上游 FunASR 与 ModelScope 开源协议(Apache-2.0 / ModelScope 社区许可)。商业集成或二次分发时,请保留上游版权声明、许可说明及原始出处链接。
文件校验清单 (SHA-256)
下载后请核对文件大小和 SHA256 校验和。
Paraformer-Large-Contextual ASR Model (ONNX Quantized) (English Documentation)
An end-to-end non-autoregressive speech recognition model with contextual hotword biasing for Mandarin Chinese ASR. This directory contains quantized ONNX weights, embedding biaser network, segmentation dictionary, vocabulary mappings, and configuration files.
Model Information
- Model Name: Paraformer-Large-Contextual ASR (NAT-ZH-CN-16k-Common-Vocab8404)
- Task: Offline Automatic Speech Recognition (ASR) with dynamic hotword enhancement
- Architecture: Non-Autoregressive Transformer + Contextual Biasing Network
- Language: Mandarin Chinese (zh-cn)
- Vocabulary Size: 8,404 Chinese characters and special tokens
- Key Feature: Supports runtime hotword injection (e.g., entity names, specialized terms, branded keywords) to substantially improve keyword recall and precision.
- Primary Files:
model_quant.onnx: INT8 statically quantized acoustic inference backbonemodel_eb.onnx: Embedding biaser network for contextual hotword injectionseg_dict: Word segmentation dictionary for hotword tokenizationtokens.json: Character and token mapping dictionaryconfig.yaml: Model configuration and biaser decoding parametersam.mvn: Cepstral mean and variance normalization (CMVN) statistics
Hardware Footprint & Sizing (CPU)
- Memory Footprint (RAM): ~1.1 GB to 1.6 GB (Including biaser model and dictionary cache, peak ~1.4 GB)
- Expected RTF (Real-Time Factor): RTF of ~0.12 - 0.22 on typical 8-core CPUs with standard hotword lists (processing a 10s audio clip takes ~1.2s - 2.2s)
- Recommended System: 4+ CPU cores, 8+ GB RAM
Audio Input Specifications
- Sampling Rate: 16000 Hz (16 kHz)
- Channels: Single channel (Mono)
- Format: 16-bit PCM or normalized float array within [-1.0, 1.0]
- Feature Extraction: 80-dimensional log Mel-filterbank features with CMVN
Recommended Pipeline Integration
- Upstream Preprocessing: Recommended pairing with
speech_fsmn_vad_zh-cn-16k-common-onnxfor voice activity detection and sentence segmentation. - Downstream Postprocessing: Paraformer outputs raw text without punctuation. It is recommended to cascade a punctuation restoration model (such as
punc_ct-transformer_zh-cn-common-vocab272727-onnx). - Hotword Usage: Hotword inputs are typically supplied as whitespace- or newline-separated lists. When no hotwords are provided, performance is equivalent to standard Paraformer-Large.
Upstream Sources
- Upstream Project: `alibaba-damo-academy/FunASR`
- ModelScope Repository: `damo/speech_paraformer-large-contextual_asr_nat-zh-cn-16k-common-vocab8404-onnx`
- Supported Inference Engines: ONNX Runtime, sherpa-onnx, FunASR Runtime
This directory provides a quantized ONNX inference distribution derived from upstream checkpoints. All copyrights belong to the original authors and Alibaba DAMO Academy / Tongyi Lab.
Licensing and Redistribution
Governed by the upstream FunASR and ModelScope licensing terms (Apache-2.0 / ModelScope Community License). Please retain upstream copyright statements, notices, and source links in downstream redistributions.
File Verification Manifest (SHA-256)
Verify file sizes and SHA-256 checksums after downloading.
