CoolFace
Datasetpublic

hotchpotch/wikipedia-multilingual-ir-pairs

wikipedia-multilingual-ir-pairs This dataset is designed for multilingual information retrieval training. Compared with raw Wikipedia paragraph dumps, it provides cleaner and more practical supervision by pairing title/section-driven queries with relevant paragraph-level documents and applying rule-based filtering to remove low-value sections and noisy fragments. It is intended for IR model training, including contrastive learning and retrieval/reranking objectives.… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-ir-pairs.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes913downloads
18 commits on main
d8ecadc4mo ago

Clarify Wikipedia license terms in dataset card

hotchpotch
452fe054mo ago

Remove construction warning from dataset card

hotchpotch
f6eff238mo ago

Update README body while preserving existing metadata

hotchpotch
e6423678mo ago

Add dataset build script referenced in README

hotchpotch
6f634028mo ago

Update README body while preserving existing metadata

hotchpotch
68682848mo ago

Update README body while preserving existing metadata

hotchpotch
cbe271d8mo ago

Upload dataset

hotchpotch
c89d4698mo ago

Upload dataset

hotchpotch
622509d8mo ago

Upload dataset

hotchpotch
f6c9de38mo ago

Upload dataset

hotchpotch
1dc896a8mo ago

Upload dataset

hotchpotch
1fb964e8mo ago

Upload dataset

hotchpotch
c5c0ed18mo ago

Upload dataset

hotchpotch
f7617ca8mo ago

Upload dataset

hotchpotch
74f7c3b8mo ago

Upload dataset

hotchpotch
fb2e2998mo ago

Upload dataset

hotchpotch
855c6118mo ago

Upload dataset

hotchpotch
7e11dad8mo ago

initial commit

hotchpotch