ChenZ95/text2vec-base-chinese
091
1---2license: apache-2.03pipeline_tag: sentence-similarity4tags:5- Sentence Transformers6- sentence-similarity7- sentence-transformers8datasets:9- shibing624/nli_zh10language:11- zh12library_name: sentence-transformers13---14 15 16# shibing624/text2vec-base-chinese17This is a CoSENT(Cosine Sentence) model: shibing624/text2vec-base-chinese.18 19It maps sentences to a 768 dimensional dense vector space and can be used for tasks 20like sentence embeddings, text matching or semantic search.21 22 23## Evaluation24For an automated evaluation of this model, see the *Evaluation Benchmark*: [text2vec](https://github.com/shibing624/text2vec)25 26- chinese text matching task:27 28| Arch | BaseModel | Model | ATEC | BQ | LCQMC | PAWSX | STS-B | SOHU-dd | SOHU-dc | Avg | QPS |29|:-----------|:----------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------|:-----:|:-----:|:-----:|:-----:|:-----:|:-------:|:-------:|:---------:|:-----:|30| Word2Vec | word2vec | [w2v-light-tencent-chinese](https://ai.tencent.com/ailab/nlp/en/download.html) | 20.00 | 31.49 | 59.46 | 2.57 | 55.78 | 55.04 | 20.70 | 35.03 | 23769 |31| SBERT | xlm-roberta-base | [sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2](https://huggingface.co/sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2) | 18.42 | 38.52 | 63.96 | 10.14 | 78.90 | 63.01 | 52.28 | 46.46 | 3138 |32| Instructor | hfl/chinese-roberta-wwm-ext | [moka-ai/m3e-base](https://huggingface.co/moka-ai/m3e-base) | 41.27 | 63.81 | 74.87 | 12.20 | 76.96 | 75.83 | 60.55 | 57.93 | 2980 |33| CoSENT | hfl/chinese-macbert-base | [shibing624/text2vec-base-chinese](https://huggingface.co/shibing624/text2vec-base-chinese) | 31.93 | 42.67 | 70.16 | 17.21 | 79.30 | 70.27 | 50.42 | 51.61 | 3008 |34| CoSENT | hfl/chinese-lert-large | [GanymedeNil/text2vec-large-chinese](https://huggingface.co/GanymedeNil/text2vec-large-chinese) | 32.61 | 44.59 | 69.30 | 14.51 | 79.44 | 73.01 | 59.04 | 53.12 | 2092 |35| CoSENT | nghuyong/ernie-3.0-base-zh | [shibing624/text2vec-base-chinese-sentence](https://huggingface.co/shibing624/text2vec-base-chinese-sentence) | 43.37 | 61.43 | 73.48 | 38.90 | 78.25 | 70.60 | 53.08 | 59.87 | 3089 |36| CoSENT | nghuyong/ernie-3.0-base-zh | [shibing624/text2vec-base-chinese-paraphrase](https://huggingface.co/shibing624/text2vec-base-chinese-paraphrase) | 44.89 | 63.58 | 74.24 | 40.90 | 78.93 | 76.70 | 63.30 | 63.08 | 3066 |37| CoSENT | sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 | [shibing624/text2vec-base-multilingual](https://huggingface.co/shibing624/text2vec-base-multilingual) | 32.39 | 50.33 | 65.64 | 32.56 | 74.45 | 68.88 | 51.17 | 53.67 | 4004 |38 39 40说明:41- 结果评测指标:spearman系数42- `shibing624/text2vec-base-chinese`模型,是用CoSENT方法训练,基于`hfl/chinese-macbert-base`在中文STS-B数据训练得到,并在中文STS-B测试集评估达到较好效果,运行[examples/training_sup_text_matching_model.py](https://github.com/shibing624/text2vec/blob/master/examples/training_sup_text_matching_model.py)代码可训练模型,模型文件已经上传HF model hub,中文通用语义匹配任务推荐使用43- `shibing624/text2vec-base-chinese-sentence`模型,是用CoSENT方法训练,基于`nghuyong/ernie-3.0-base-zh`用人工挑选后的中文STS数据集[shibing624/nli-zh-all/text2vec-base-chinese-sentence-dataset](https://huggingface.co/datasets/shibing624/nli-zh-all/tree/main/text2vec-base-chinese-sentence-dataset)训练得到,并在中文各NLI测试集评估达到较好效果,运行[examples/training_sup_text_matching_model_jsonl_data.py](https://github.com/shibing624/text2vec/blob/master/examples/training_sup_text_matching_model_jsonl_data.py)代码可训练模型,模型文件已经上传HF model hub,中文s2s(句子vs句子)语义匹配任务推荐使用44- `shibing624/text2vec-base-chinese-paraphrase`模型,是用CoSENT方法训练,基于`nghuyong/ernie-3.0-base-zh`用人工挑选后的中文STS数据集[shibing624/nli-zh-all/text2vec-base-chinese-paraphrase-dataset](https://huggingface.co/datasets/shibing624/nli-zh-all/tree/main/text2vec-base-chinese-paraphrase-dataset),数据集相对于[shibing624/nli-zh-all/text2vec-base-chinese-sentence-dataset](https://huggingface.co/datasets/shibing624/nli-zh-all/tree/main/text2vec-base-chinese-sentence-dataset)加入了s2p(sentence to paraphrase)数据,强化了其长文本的表征能力,并在中文各NLI测试集评估达到SOTA,运行[examples/training_sup_text_matching_model_jsonl_data.py](https://github.com/shibing624/text2vec/blob/master/examples/training_sup_text_matching_model_jsonl_data.py)代码可训练模型,模型文件已经上传HF model hub,中文s2p(句子vs段落)语义匹配任务推荐使用45- `sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2`模型是用SBERT训练,是`paraphrase-MiniLM-L12-v2`模型的多语言版本,支持中文、英文等46- `w2v-light-tencent-chinese`是腾讯词向量的Word2Vec模型,CPU加载使用,适用于中文字面匹配任务和缺少数据的冷启动情况47 48## Usage (text2vec)49Using this model becomes easy when you have [text2vec](https://github.com/shibing624/text2vec) installed:50 51```52pip install -U text2vec53```54 55Then you can use the model like this:56 57```python58from text2vec import SentenceModel59sentences = ['如何更换花呗绑定银行卡', '花呗更改绑定银行卡']60 61model = SentenceModel('shibing624/text2vec-base-chinese')62embeddings = model.encode(sentences)63print(embeddings)64```65 66## Usage (HuggingFace Transformers)67Without [text2vec](https://github.com/shibing624/text2vec), you can use the model like this: 68 69First, you pass your input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings.70 71Install transformers:72```73pip install transformers74```75 76Then load model and predict:77```python78from transformers import BertTokenizer, BertModel79import torch80 81# Mean Pooling - Take attention mask into account for correct averaging82def mean_pooling(model_output, attention_mask):83 token_embeddings = model_output[0] # First element of model_output contains all token embeddings84 input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()85 return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(input_mask_expanded.sum(1), min=1e-9)86 87# Load model from HuggingFace Hub88tokenizer = BertTokenizer.from_pretrained('shibing624/text2vec-base-chinese')89model = BertModel.from_pretrained('shibing624/text2vec-base-chinese')90sentences = ['如何更换花呗绑定银行卡', '花呗更改绑定银行卡']91# Tokenize sentences92encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')93 94# Compute token embeddings95with torch.no_grad():96 model_output = model(**encoded_input)97# Perform pooling. In this case, mean pooling.98sentence_embeddings = mean_pooling(model_output, encoded_input['attention_mask'])99print("Sentence embeddings:")100print(sentence_embeddings)101```102 103## Usage (sentence-transformers)104[sentence-transformers](https://github.com/UKPLab/sentence-transformers) is a popular library to compute dense vector representations for sentences.105 106Install sentence-transformers:107```108pip install -U sentence-transformers109```110 111Then load model and predict:112 113```python114from sentence_transformers import SentenceTransformer115 116m = SentenceTransformer("shibing624/text2vec-base-chinese")117sentences = ['如何更换花呗绑定银行卡', '花呗更改绑定银行卡']118 119sentence_embeddings = m.encode(sentences)120print("Sentence embeddings:")121print(sentence_embeddings)122```123 124## Model speed up125 126 127| Model | ATEC | BQ | LCQMC | PAWSX | STSB |128|------------------------------------------------------------------------------------------------------------------------------|-------------------|-------------------|------------------|------------------|------------------|129| shibing624/text2vec-base-chinese (fp32, baseline) | 0.31928 | 0.42672 | 0.70157 | 0.17214 | 0.79296 |130| shibing624/text2vec-base-chinese (onnx-O4, [#29](https://huggingface.co/shibing624/text2vec-base-chinese/discussions/29)) | 0.31928 | 0.42672 | 0.70157 | 0.17214 | 0.79296 |131| shibing624/text2vec-base-chinese (ov, [#27](https://huggingface.co/shibing624/text2vec-base-chinese/discussions/27)) | 0.31928 | 0.42672 | 0.70157 | 0.17214 | 0.79296 |132| shibing624/text2vec-base-chinese (ov-qint8, [#30](https://huggingface.co/shibing624/text2vec-base-chinese/discussions/30)) | 0.30778 (-3.60%) | 0.43474 (+1.88%) | 0.69620 (-0.77%) | 0.16662 (-3.20%) | 0.79396 (+0.13%) |133 134In short:1351. ✅ shibing624/text2vec-base-chinese (onnx-O4), ONNX Optimized to [O4](https://huggingface.co/docs/optimum/en/onnxruntime/usage_guides/optimization) does not reduce performance, but gives a [~2x speedup](https://sbert.net/docs/sentence_transformer/usage/efficiency.html#benchmarks) on GPU.1362. ✅ shibing624/text2vec-base-chinese (ov), OpenVINO does not reduce performance, but gives a 1.12x speedup on CPU.1373. 🟡 shibing624/text2vec-base-chinese (ov-qint8), int8 quantization with OV incurs a small performance hit on some tasks, and a tiny performance gain on others, when quantizing with [Chinese STSB](https://huggingface.co/datasets/PhilipMay/stsb_multi_mt). Additionally, it results in a [4.78x speedup](https://sbert.net/docs/sentence_transformer/usage/efficiency.html#benchmarks) on CPU.138 139- usage: shibing624/text2vec-base-chinese (onnx-O4), for gpu140```python141from sentence_transformers import SentenceTransformer142 143model = SentenceTransformer(144 "shibing624/text2vec-base-chinese",145 backend="onnx",146 model_kwargs={"file_name": "model_O4.onnx"},147)148embeddings = model.encode(["如何更换花呗绑定银行卡", "花呗更改绑定银行卡", "你是谁"])149print(embeddings.shape)150similarities = model.similarity(embeddings, embeddings)151print(similarities)152```153 154 155- usage: shibing624/text2vec-base-chinese (ov), for cpu156```python157# pip install 'optimum[openvino]'158 159from sentence_transformers import SentenceTransformer160 161model = SentenceTransformer(162 "shibing624/text2vec-base-chinese",163 backend="openvino",164)165 166embeddings = model.encode(["如何更换花呗绑定银行卡", "花呗更改绑定银行卡", "你是谁"])167print(embeddings.shape)168similarities = model.similarity(embeddings, embeddings)169print(similarities)170```171 172- usage: shibing624/text2vec-base-chinese (ov-qint8), for cpu173```python174# pip install optimum175from sentence_transformers import SentenceTransformer176 177model = SentenceTransformer(178 "shibing624/text2vec-base-chinese",179 backend="onnx",180 model_kwargs={"file_name": "model_qint8_avx512_vnni.onnx"},181)182embeddings = model.encode(["如何更换花呗绑定银行卡", "花呗更改绑定银行卡", "你是谁"])183print(embeddings.shape)184similarities = model.similarity(embeddings, embeddings)185print(similarities)186```187 188 189## Full Model Architecture190```191CoSENT(192 (0): Transformer({'max_seq_length': 128, 'do_lower_case': False}) with Transformer model: BertModel 193 (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_mean_tokens': True})194)195```196 197## Intended uses198 199Our model is intented to be used as a sentence and short paragraph encoder. Given an input text, it ouptuts a vector which captures 200the semantic information. The sentence vector may be used for information retrieval, clustering or sentence similarity tasks.201 202By default, input text longer than 256 word pieces is truncated.203 204 205## Training procedure206 207### Pre-training 208 209We use the pretrained [`hfl/chinese-macbert-base`](https://huggingface.co/hfl/chinese-macbert-base) model. 210Please refer to the model card for more detailed information about the pre-training procedure.211 212### Fine-tuning 213 214We fine-tune the model using a contrastive objective. Formally, we compute the cosine similarity from each 215possible sentence pairs from the batch.216We then apply the rank loss by comparing with true pairs and false pairs.217 218#### Hyper parameters219 220- training dataset: https://huggingface.co/datasets/shibing624/nli_zh221- max_seq_length: 128222- best epoch: 5223- sentence embedding dim: 768224 225 226 227## Citing & Authors228This model was trained by [text2vec](https://github.com/shibing624/text2vec). 229 230If you find this model helpful, feel free to cite:231```bibtex 232@software{text2vec,233 author = {Xu Ming},234 title = {text2vec: A Tool for Text to Vector},235 year = {2022},236 url = {https://github.com/shibing624/text2vec},237}238```