ranger810/sherpa-onnx-punct-ct-transformer-zh-en-vocab272727-2024-04-12-int8
043
sherpa-onnx-punct-ct-transformer-zh-en-vocab272727-2024-04-12-int8
A Chinese/English punctuation restoration model based on the Alibaba DAMO Academy CT-Transformer architecture, exported to ONNX INT8 quantized format using the sherpa-onnx framework.
Model Overview
File List
Usage
Install sherpa-onnx
pip install sherpa-onnxPython Example
import sherpa_onnx
config = sherpa_onnx.OfflinePunctuationConfig(
model=sherpa_onnx.OfflinePunctuationModelConfig(
ct_transformer="./model.int8.onnx",
),
tokens="./tokens.json",
)
punct = sherpa_onnx.OfflinePunctuation(config)
text = "我们都是中国人我爱中国"
result = punct.add_punctuation(text)
print(result)
# Output: 我们都是中国人,我爱中国。Run Test Script
python test.pyModel Origin
This model is converted from the following original model by the sherpa-onnx project:
- Original Model: ModelScope - iic/punc_ct-transformer_zh-cn-common-vocab272727-pytorch
- Conversion: Exported via sherpa-onnx export scripts with INT8 quantization
- Original Paper: CT-Transformer: Controllable Time-delay Transformer for Real-Time Punctuation Prediction and Disfluency Detection
License
This model is licensed under the Apache License 2.0.
The original model is copyright of Alibaba DAMO Academy and is also released under the Apache 2.0 License.
