liboaccn/nmt-parallel-corpus
Neural Machine Translation parallel corpora Introduction We use OpusTools to extract resources from the OPUS project, a renowned platform for parallel corpora, and create a multilingual dataset. Specifically, we collect the parallel corpora from prominent projects within OPUS, including NLLB, CCMatrix, and OpenSubtitles. This comprehensive data collection process results in a corpus of more than 3T, covering 60 languages and over 1900 language pairs.… See the full description on the dataset page: https://huggingface.co/datasets/liboaccn/nmt-parallel-corpus.
This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.
