Proyag/paracrawl_context
Dataset Card for ParaCrawl_Context This is a dataset for document-level machine translation introduced in the ACL 2024 paper Document-Level Machine Translation with Large-Scale Public Parallel Data. It is a dataset consisting of parallel sentence pairs from the ParaCrawl dataset along with corresponding preceding context extracted from the webpages the sentences were crawled from. Dataset Details Dataset Description This dataset adds… See the full description on the dataset page: https://huggingface.co/datasets/Proyag/paracrawl_context.
This repository belongs to Proyag on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
