CoolFace
Datasetpublic

sentence-transformers/parallel-sentences-europarl

Dataset Card for Parallel Sentences - Europarl This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Europarl dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-europarl.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes1.8kdownloads
Dataset Card

Dataset Card for Parallel Sentences - Europarl

This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Europarl dataset.

Related Datasets

The following datasets are also a part of the Parallel Sentences collection:

These datasets can be used to train multilingual sentence embedding models. For more information, see sbert.net - Multilingual Models.

Dataset Subsets

all subset

  • Columns: "english", "non_english"
  • Column types: str, str
  • Examples:
python
    {
      "english": "Membership of Parliament: see Minutes",
  	  "non_english": "Състав на Парламента: вж. протоколи"
    }
  • Collection strategy: Combining all other subsets from this dataset.
  • Deduplified: No

en-... subsets

  • Columns: "english", "non_english"
  • Column types: str, str
  • Examples:
python
    {
      "english": "Resumption of the session",
      "non_english": "Reanudación del período de sesiones"
    }
  • Collection strategy: Processing the raw data from parallel-sentences and formatting it in Parquet, followed by deduplication.
  • Deduplified: Yes