CoolFace
Datasetpublic

Helsinki-NLP/nemotron-cc-translated

Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the high-quality subset. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 156,431,999 documents with over 70 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 2.4 trillion tokens and the translated documents are aligned across all languages. v1.1… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/nemotron-cc-translated.

sourceHugging Facecc0-1.0updated 5mo agoView on Hugging Face
5likes56kdownloads
settings

This repository belongs to Helsinki-NLP on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenemotron-cc-translated
visibilitypublic
licencecc0-1.0
gatedno
ownerHelsinki-NLP
Account settings