CoolFace
Datasetpublic

wecover/OPUS

Collection of OPUS Corpus from https://opus.nlpl.eu has been collected. The following corpora have been included: UNPC GlobalVoices TED2020 News-Commentary WikiMatrix Tatoeba Europarl OpenSubtitles 25,000 samples (randomly sampled within the first 100,000 samples) per language pair of each corpus were collected, with no modification of data. Licenses OPUS @inproceedings{tiedemann2012parallel, title={Parallel data, tools and interfaces in OPUS.}… See the full description on the dataset page: https://huggingface.co/datasets/wecover/OPUS.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes5.2kdownloads
21 commits on main
21c82512y ago

Update README.md

0601p
fc513852y ago

Update README.md

0601p
b29ac202y ago

Delete READMD.md

0601p
2b553e92y ago

Create README.md

0601p
2cb8c453y ago

fix punctuation marks issues

0601p
ea864983y ago

remove data that overlaps with mteb

0601p
eb797e23y ago

Update READMD.md

0601p
f139cc33y ago

fix config

0601p
926b91a3y ago

debug2

0601p
78a49cd3y ago

debug

0601p
d9b3d483y ago

remove langs

0601p
bc188413y ago

update readme

0601p
76510c23y ago

OpenSubtitles added

0601p
49ce64a3y ago

WikiMatrix added

0601p
8b43d373y ago

TED2020 added

0601p
950b42c3y ago

Europarl added

0601p
1d76f773y ago

GlobalVoices added

0601p
a956ef83y ago

News-Commentary added

0601p
ec177f03y ago

UNPC added

0601p
005e4713y ago

Tatoeba added

0601p
5c1172d3y ago

initial commit

0601p