HPLT/DocHPLT
DocHPLT: A Massively Multilingual Document-Level Translation Dataset Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, long-context modeling for global communities, we create DocHPLT, the largest publicly available document-level translation dataset to date. It contains 124 million aligned document pairs across… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/DocHPLT.
2015k
../
train-00000-of-00025.parquetdownload
train-00001-of-00025.parquetdownload
train-00002-of-00025.parquetdownload
train-00003-of-00025.parquetdownload
train-00004-of-00025.parquetdownload
train-00005-of-00025.parquetdownload
train-00006-of-00025.parquetdownload
train-00007-of-00025.parquetdownload
train-00008-of-00025.parquetdownload
train-00009-of-00025.parquetdownload
train-00010-of-00025.parquetdownload
train-00011-of-00025.parquetdownload
train-00012-of-00025.parquetdownload
train-00013-of-00025.parquetdownload
train-00014-of-00025.parquetdownload
train-00015-of-00025.parquetdownload
train-00016-of-00025.parquetdownload
train-00017-of-00025.parquetdownload
train-00018-of-00025.parquetdownload
train-00019-of-00025.parquetdownload
train-00020-of-00025.parquetdownload
train-00021-of-00025.parquetdownload
train-00022-of-00025.parquetdownload
train-00023-of-00025.parquetdownload
train-00024-of-00025.parquetdownload
