CoolFace
Datasetpublic

wecover/OPUS

Collection of OPUS Corpus from https://opus.nlpl.eu has been collected. The following corpora have been included: UNPC GlobalVoices TED2020 News-Commentary WikiMatrix Tatoeba Europarl OpenSubtitles 25,000 samples (randomly sampled within the first 100,000 samples) per language pair of each corpus were collected, with no modification of data. Licenses OPUS @inproceedings{tiedemann2012parallel, title={Parallel data, tools and interfaces in OPUS.}… See the full description on the dataset page: https://huggingface.co/datasets/wecover/OPUS.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes5.2kdownloads

wecover/OPUS · main · files are served by the source, never re-hosted here

wecover/OPUS · CoolFace