wecover/OPUS
Collection of OPUS Corpus from https://opus.nlpl.eu has been collected. The following corpora have been included: UNPC GlobalVoices TED2020 News-Commentary WikiMatrix Tatoeba Europarl OpenSubtitles 25,000 samples (randomly sampled within the first 100,000 samples) per language pair of each corpus were collected, with no modification of data. Licenses OPUS @inproceedings{tiedemann2012parallel, title={Parallel data, tools and interfaces in OPUS.}… See the full description on the dataset page: https://huggingface.co/datasets/wecover/OPUS.
Update README.md
Update README.md
Delete READMD.md
Create README.md
fix punctuation marks issues
remove data that overlaps with mteb
Update READMD.md
fix config
debug2
debug
remove langs
update readme
OpenSubtitles added
WikiMatrix added
TED2020 added
Europarl added
GlobalVoices added
News-Commentary added
UNPC added
Tatoeba added
initial commit
