CoolFace
Datasetpublic

silma-ai/silma-arabic-english-sts-dataset-v1.0

SILMA STS Arabic/English Dataset - v1.0 Overview The SILMA STS Arabic/English Dataset - v1.0 is a dataset designed for training and evaluating sentence embeddings for Arabic and English tasks. It consists of five different splits that cover monolingual and multilingual sentence pairs, with human-annotated similarity scores. The dataset includes both Arabic-to-Arabic and English-to-English pairs, as well as cross-lingual Arabic-English pairs, making it a valuable… See the full description on the dataset page: https://huggingface.co/datasets/silma-ai/silma-arabic-english-sts-dataset-v1.0.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
3likes40downloads
Dataset Card

SILMA STS Arabic/English Dataset - v1.0

Overview

The SILMA STS Arabic/English Dataset - v1.0 is a dataset designed for training and evaluating sentence embeddings for Arabic and English tasks. It consists of five different splits that cover monolingual and multilingual sentence pairs, with human-annotated similarity scores. The dataset includes both Arabic-to-Arabic and English-to-English pairs, as well as cross-lingual Arabic-English pairs, making it a valuable resource for multilingual and cross-lingual semantic similarity tasks.

Dataset Structure

The dataset is divided into five splits, each containing sentence pairs and similarity scores.

Split 1: ar_ar

  • —Description: Contains Arabic-to-Arabic sentence pairs with similarity scores.
  • —Size: 11,512 examples
  • —JSON Sample:
json
  {
    "sentence1": "رجلين يلعبان الشطرنج",
    "sentence2": "ثلاثة رجال يلعبون الشطرنج",
    "score": 0.52,
    "langs": "ar-ar"
  }

### Split 2: en_en
- **Description:** Contains English-to-English sentence pairs with similarity scores.
- **Size:** 11,512 examples
- **JSON Sample:**

{ "sentence1": "A plane is taking off.", "sentence2": "An air plane is taking off.", "score": 1.0 }

Split 3: multilingual_1

  • —Description: Contains sentence pairs from both Arabic and English, with similarity scores. The sentences are aligned cross-lingually.
  • —Size: 5,756 examples
  • —JSON Sample:
json
  {
    "sentence1": "The man is playing the guitar. | الرجل يعزف على الغيتار",
    "sentence2": "The man is playing the piano. | الرجل يعزف على البيانو",
    "score": 0.32
  }

### Split 4: multilingual_2
- **Description:** Similar to Split 3, but with reversed language pairs.
- **Size:** 5,756 examples
- **JSON Sample:**

{ "sentence1": "رجل يدخن | A man is smoking.", "sentence2": "رجل يتزلج | A man is skating.", "score": 0.1 }

Column Descriptions

Each split in the dataset contains the following columns:

  • —sentence1: The first sentence in the pair. It can be in Arabic or English depending on the split.
  • —sentence2: The second sentence in the pair. It can also be in Arabic or English depending on the split.
  • —score: A floating-point number between 0 and 1 representing the semantic similarity between the two sentences, where 1 indicates maximum similarity.
  • —langs: Indicates the language pair of the sentences. The possible values are:
  • —ar-ar (Arabic-Arabic)
  • —en-en (English-English)
  • —Multilingual-1 (Multilingual, English-Arabic)
  • —Multilingual-2 (Multilingual, Arabic-English)

Use Cases

The SILMA STS Arabic/English Dataset - v1.0 can be used in various NLP tasks, including but not limited to:

  1. 1.Sentence Embedding Training: The dataset is well-suited for training models that generate sentence embeddings, enabling effective comparison of sentence-level semantics in both Arabic and English.
  2. 2.Multilingual and Cross-Lingual STS: This dataset can be used for evaluating the performance of multilingual and cross-lingual sentence transformers, as it includes both monolingual and multilingual sentence pairs.
  3. 3.Semantic Similarity Tasks: The dataset can be utilized in semantic similarity benchmarks, particularly for Arabic and English language pairs.
  4. 4.Cross-Lingual Transfer Learning: The multilingual sentence pairs provide a good opportunity for training models in cross-lingual transfer learning, where knowledge from one language can be transferred to another.

This dataset is a useful resource for researchers and developers working on NLP tasks that involve sentence semantics across different languages, especially for Arabic and English.