CoolFace
Modelpublic

ranger810/sherpa-onnx-punct-ct-transformer-zh-en-vocab272727-2024-04-12-int8

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes43downloads
Model Card

sherpa-onnx-punct-ct-transformer-zh-en-vocab272727-2024-04-12-int8

A Chinese/English punctuation restoration model based on the Alibaba DAMO Academy CT-Transformer architecture, exported to ONNX INT8 quantized format using the sherpa-onnx framework.

Model Overview

AttributeDetails
Original Modeliic/punc_ct-transformer_zh-cn-common-vocab272727-pytorch
Model FormatONNX INT8 Quantized
Supported LanguagesChinese, English
Vocabulary Size272,727
Frameworksherpa-onnx
LicenseApache 2.0

File List

FilenameDescription
model.int8.onnxINT8 quantized ONNX model file
tokens.jsonVocabulary file
config.yamlModel configuration file
test.pyTest script
show-model-input-output.pyScript to inspect model input/output structure
add-model-metadata.pyScript to add model metadata

Usage

Install sherpa-onnx

bash
pip install sherpa-onnx

Python Example

python
import sherpa_onnx

config = sherpa_onnx.OfflinePunctuationConfig(
    model=sherpa_onnx.OfflinePunctuationModelConfig(
        ct_transformer="./model.int8.onnx",
    ),
    tokens="./tokens.json",
)

punct = sherpa_onnx.OfflinePunctuation(config)

text = "我们都是中国人我爱中国"
result = punct.add_punctuation(text)
print(result)
# Output: 我们都是中国人,我爱中国。

Run Test Script

bash
python test.py

Model Origin

This model is converted from the following original model by the sherpa-onnx project:

License

This model is licensed under the Apache License 2.0.

The original model is copyright of Alibaba DAMO Academy and is also released under the Apache 2.0 License.

Related Links