clips/beir-nl-mmarco
Dataset Card for BEIR-NL Benchmark BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB). Source Data mMarco repository on GitHub. Additional Information Licensing Information This dataset is licensed… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-mmarco.
Dataset Card for BEIR-NL Benchmark
BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB).
Source Data
mMarco repository on GitHub.
Additional Information
Licensing Information
This dataset is licensed under the Apache license 2.0..
Citation Information
We use this dataset in our work:
@misc{banar2024beirnlzeroshotinformationretrieval,
title={BEIR-NL: Zero-shot Information Retrieval Benchmark for the Dutch Language},
author={Nikolay Banar and Ehsan Lotfi and Walter Daelemans},
year={2024},
eprint={2412.08329},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.08329},
}The original source of data:
@article{DBLP:journals/corr/abs-2108-13897,
author = {Luiz Bonifacio and
Israel Campiotti and
Roberto de Alencar Lotufo and
Rodrigo Frassetto Nogueira},
title = {mMARCO: {A} Multilingual Version of {MS} {MARCO} Passage Ranking Dataset},
journal = {CoRR},
volume = {abs/2108.13897},
year = {2021},
url = {https://arxiv.org/abs/2108.13897},
eprinttype = {arXiv},
eprint = {2108.13897},
timestamp = {Mon, 20 Mar 2023 15:35:34 +0100},
biburl = {https://dblp.org/rec/journals/corr/abs-2108-13897.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}