CoolFace
Datasetpublic

sentence-transformers/stsb

Dataset Card for STSB The Semantic Textual Similarity Benchmark (Cer et al., 2017) is a collection of sentence pairs drawn from news headlines, video and image captions, and natural language inference data. Each pair is human-annotated with a similarity score from 1 to 5. However, for this variant, the similarity scores are normalized to between 0 and 1. Dataset Details Columns: "sentence1", "sentence2", "score" Column types: str, str, float Examples:{… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/stsb.

sourceHugging Faceupdated 2y agoView on Hugging Face
26likes16kdownloads
Dataset Card

Dataset Card for STSB

The Semantic Textual Similarity Benchmark (Cer et al., 2017) is a collection of sentence pairs drawn from news headlines, video and image captions, and natural language inference data. Each pair is human-annotated with a similarity score from 1 to 5. However, for this variant, the similarity scores are normalized to between 0 and 1.

Dataset Details

  • Columns: "sentence1", "sentence2", "score"
  • Column types: str, str, float
  • Examples:
python
    {
      'sentence1': 'A man is playing a large flute.',
      'sentence2': 'A man is playing a flute.',
      'score': 0.76,
    }
  • Collection strategy: Reading the sentences and score from STSB dataset and dividing the score by 5.
  • Deduplified: No