sentence-transformers/parallel-sentences-wikimatrix
Dataset Card for Parallel Sentences - WikiMatrix This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the WikiMatrix dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-wikimatrix.
Dataset Card for Parallel Sentences - WikiMatrix
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the WikiMatrix dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
- parallel-sentences-europarl
- parallel-sentences-global-voices
- parallel-sentences-muse
- parallel-sentences-jw300
- parallel-sentences-news-commentary
- parallel-sentences-opensubtitles
- parallel-sentences-talks
- parallel-sentences-tatoeba
- parallel-sentences-wikimatrix
- parallel-sentences-wikititles
- parallel-sentences-ccmatrix
These datasets can be used to train multilingual sentence embedding models. For more information, see sbert.net - Multilingual Models.
Dataset Subsets
all subset
- Columns: "english", "non_english"
- Column types:
str,str - Examples:
{
"english": "I will go down with you into Egypt, and I will also surely bring you up again."",
"non_english": "Аз ще бъда с тебе в Египет и Аз ще те изведа назад.""
}- Collection strategy: Combining all other subsets from this dataset.
- Deduplified: No
en-... subsets
- Columns: "english", "non_english"
- Column types:
str,str - Examples:
{
"english": "By Him Who is the Lord of mankind!",
"non_english": "しかしながら人主の患はまた人を信ぜざるにもある。"
}- Collection strategy: Processing the raw data from parallel-sentences and formatting it in Parquet, followed by deduplication.
- Deduplified: Yes
