CoolFace
Datasetpublicgated

liboaccn/nmt-parallel-corpus

Neural Machine Translation parallel corpora Introduction We use OpusTools to extract resources from the OPUS project, a renowned platform for parallel corpora, and create a multilingual dataset. Specifically, we collect the parallel corpora from prominent projects within OPUS, including NLLB, CCMatrix, and OpenSubtitles. This comprehensive data collection process results in a corpus of more than 3T, covering 60 languages and over 1900 language pairs.… See the full description on the dataset page: https://huggingface.co/datasets/liboaccn/nmt-parallel-corpus.

sourceHugging Facecc-by-nc-4.0updated 1y agoView on Hugging Face
24likes28downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.