CoolFace
Datasetpublic

HPLT/DocHPLT

DocHPLT: A Massively Multilingual Document-Level Translation Dataset Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, long-context modeling for global communities, we create DocHPLT, the largest publicly available document-level translation dataset to date. It contains 124 million aligned document pairs across… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/DocHPLT.

sourceHugging Facecc0-1.0updated 8mo agoView on Hugging Face
20likes15kdownloads
../
filetrain-00000-of-00010.parquet230.6 MBdownload
filetrain-00001-of-00010.parquet118.7 MBdownload
filetrain-00002-of-00010.parquet119.5 MBdownload
filetrain-00003-of-00010.parquet120.4 MBdownload
filetrain-00004-of-00010.parquet166.6 MBdownload
filetrain-00005-of-00010.parquet165.3 MBdownload
filetrain-00006-of-00010.parquet160.8 MBdownload
filetrain-00007-of-00010.parquet159.2 MBdownload
filetrain-00008-of-00010.parquet126.8 MBdownload
filetrain-00009-of-00010.parquet141.7 MBdownload

HPLT/DocHPLT · main · files are served by the source, never re-hosted here