CoolFace
Datasetpublic

sentence-transformers/parallel-sentences-tatoeba

Dataset Card for Parallel Sentences - Tatoeba This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Tatoeba dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-tatoeba.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes3.8kdownloads
Dataset Card

Dataset Card for Parallel Sentences - Tatoeba

This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Tatoeba dataset.

Related Datasets

The following datasets are also a part of the Parallel Sentences collection:

These datasets can be used to train multilingual sentence embedding models. For more information, see sbert.net - Multilingual Models.

Dataset Subsets

all subset

  • Columns: "english", "non_english"
  • Column types: str, str
  • Examples:
python
    {
      "english": "I met a friend of Mary's.",
      "non_english": "मी मेरीच्या एका मैत्रिणीला भेटलो."
    }
  • Collection strategy: Combining all other subsets from this dataset.
  • Deduplified: No

en-... subsets

  • Columns: "english", "non_english"
  • Column types: str, str
  • Examples:
python
    {
      "english": "The password is "Muiriel".",
      "non_english": "Das Passwort lautet „Muiriel“."
    }
  • Collection strategy: Processing the raw data from parallel-sentences and formatting it in Parquet, followed by deduplication.
  • Deduplified: Yes