hotchpotch/wikipedia-multilingual-ir-pairs
wikipedia-multilingual-ir-pairs This dataset is designed for multilingual information retrieval training. Compared with raw Wikipedia paragraph dumps, it provides cleaner and more practical supervision by pairing title/section-driven queries with relevant paragraph-level documents and applying rule-based filtering to remove low-value sections and noisy fragments. It is intended for IR model training, including contrastive learning and retrieval/reranking objectives.… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-ir-pairs.
Clarify Wikipedia license terms in dataset card
Remove construction warning from dataset card
Update README body while preserving existing metadata
Add dataset build script referenced in README
Update README body while preserving existing metadata
Update README body while preserving existing metadata
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
initial commit
