CoolFace
Datasetpublic

omarelba/letzcross-wiki-parallel

LëtzCross Wiki Parallel Dataset Description omarelba/letzcross-wiki-parallel is a multilingual visual-document-retrieval dataset. Each row pairs a rendered PDF page with parallel questions in English, French, German, and, where available, Luxembourgish. The dataset also includes answer and provenance fields produced during dataset construction and validation. The dataset is used for language-specific late-interaction page-image retrieval training. A training… See the full description on the dataset page: https://huggingface.co/datasets/omarelba/letzcross-wiki-parallel.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes142downloads
../
filetest-00000-of-00001.parquet21.8 MBdownload
filetrain-00000-of-00003.parquet344.2 MBdownload
filetrain-00001-of-00003.parquet346.2 MBdownload
filetrain-00002-of-00003.parquet346.6 MBdownload

omarelba/letzcross-wiki-parallel · main · files are served by the source, never re-hosted here