CoolFace
Modelpublic

w831152001/TARS-S-BERT-Evaluator

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes179downloads
Model Card

~~This repository features an S-BERT model fine-tuned under an nn.Module evaluator.~~

~~By optimizing with a custom binary cross-entropy objective,~~

~~it effectively breaks the embedding anisotropy effect (cone effect) and aligns the original S-BERT representations with Word2Vec,~~

~~forcing the output logits to stretch dynamically across positive and negative spectrums.~~


~~This model yields a similarity score of 0.001155687728896737 between 'I like to read.' and 'Reading is a good habit.',~~

~~compared to a score of 0.03844790533185005 from the vanilla S-BERT.~~

~~Note that lower scores indicate higher semantic similarity.~~


Because original S-BERT has anisotropy effect (cone effect),

the similarity score of any sentence is very close and cannot be distinguished. Not to mention transferring.

~~So I first trained an nn.Module network making the subtraction of embedding vector of one sentence in S-BERT and in Word2Vec with another similar sentence the closer the better.~~

So I trained an nn.Module network making the subtraction of embedding vector of one sentence in S-BERT and in Word2Vec with another similar sentence the closer the better.

My dataset is mteb/stsbenchmark-sts.

~~Then I used the nn.Module network evaluator to fine-tune S-BERT.~~

Because the contextual logic of S-BERT will be destroyed if I find-tune it, so I keep evaluator(nn.Module network) which transfers S-BERT space to Word2Vec space instead(code as main.py).


~~from sentence_transformers import SentenceTransformer~~

~~model = SentenceTransformer("w831152001/TARS-S-BERT")~~

~~queryembedding = model.encode("I like to read.", convertto_tensor=True)~~

~~corpusembedding = model.encode("Reading is a good habit.", convertto_tensor=True)~~

~~distance = torch.mean(torch.abs(corpusembedding - queryembedding), dim=0).item()~~

~~print(distance)~~

~~sbertmodelbase = SentenceTransformer('sentence-transformers/all-MiniLM-L6-v2')~~

~~queryembeddingbase = sbertmodelbase.encode("I like to read.", converttotensor=True)~~

~~corpusembeddingbase = sbertmodelbase.encode("Reading is a good habit.", converttotensor=True)~~

~~distancebase = torch.mean(torch.abs(corpusembeddingbase - queryembedding_base), dim=0).item()~~

~~print(distance_base)~~


~~The output of running the demo.py and it explains the transferring:~~

The output of running the demo.py:

[input ticket]: 'I just bought a business-class flight ticket to Paris.'

vs. [Five-star Luxury Hotel in Paris] -> similarity: 0.8585350215435028

vs. [Budget Hostel near Tokyo Station] -> similarity: 0.14060966670513153

vs. [Airport Capsule Hotel] -> similarity: 0.6487089395523071

=> The system automatically recommends the most suitable accommodation: [Five-star Luxury Hotel in Paris] (highest similarity: 0.8585350215435028)

[input ticket]: 'My flight departs tomorrow morning at 6 AM.'

vs. [Five-star Luxury Hotel in Paris] -> similarity: 1.2763501703739166

vs. [Budget Hostel near Tokyo Station] -> similarity: 0.4070654511451721

vs. [Airport Capsule Hotel] -> similarity: 1.4093074202537537

=> The system automatically recommends the most suitable accommodation: [Airport Capsule Hotel] (highest similarity: 1.4093074202537537)