CoolFace
Datasetpublic

Helsinki-NLP/nemotron-cc-translated

Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the high-quality subset. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 156,431,999 documents with over 70 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 2.4 trillion tokens and the translated documents are aligned across all languages. v1.1… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/nemotron-cc-translated.

sourceHugging Facecc0-1.0updated 5mo agoView on Hugging Face
5likes56kdownloads
dirbos/
dirbul/
dircat/
dirces/
dirdan/
dirdeu/
direll/
direng/
direst/
direus/
dirfin/
dirfra/
dirgle/
dirglg/
dirhrv/
dirhun/
dirisl/
dirita/
dirkat/
dirlav/
dirlit/
dirmkd/
dirmlt/
dirnld/
dirnno/
dirnob/
dirpol/
dirpor/
dirron/
dirslk/
dirslv/
dirspa/
dirsqi/
dirswe/
dirtur/
dirukr/

Helsinki-NLP/nemotron-cc-translated · main · files are served by the source, never re-hosted here