omarelba/letzcross-wiki-parallel
LëtzCross Wiki Parallel Dataset Description omarelba/letzcross-wiki-parallel is a multilingual visual-document-retrieval dataset. Each row pairs a rendered PDF page with parallel questions in English, French, German, and, where available, Luxembourgish. The dataset also includes answer and provenance fields produced during dataset construction and validation. The dataset is used for language-specific late-interaction page-image retrieval training. A training… See the full description on the dataset page: https://huggingface.co/datasets/omarelba/letzcross-wiki-parallel.
0142
