CoolFace
Datasetpublic

omarelba/letzcross-wiki-parallel

LëtzCross Wiki Parallel Dataset Description omarelba/letzcross-wiki-parallel is a multilingual visual-document-retrieval dataset. Each row pairs a rendered PDF page with parallel questions in English, French, German, and, where available, Luxembourgish. The dataset also includes answer and provenance fields produced during dataset construction and validation. The dataset is used for language-specific late-interaction page-image retrieval training. A training… See the full description on the dataset page: https://huggingface.co/datasets/omarelba/letzcross-wiki-parallel.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes132downloads
3 commits on main
76d32531mo ago

Update README.md

omarelba
d1b44b28mo ago

Upload dataset

omarelba
6ef185b8mo ago

initial commit

omarelba