datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
touche2020SWE-Touch
SWE-Touch
SWE-Touch evaluates coding agents when a user edits the same workspace during an
ongoing software task. This release contains 250 validated records spanning
SWE-bench Verified, SWE-Bench Pro, and DeepSWE.
Each record includes task-critical regions, a validated Counter-Edit or text
fallback, its trigger schedule, and the user-simulator prompt identifier. The
construction and evaluation pipeline is available at
Trae1ounG/SWE-Touch.
Configurations… See the full description on the dataset page: https://huggingface.co/datasets/Trae1ounG/SWE-Touch.webis-touche2020-generated-queries
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/webis-touche2020-generated-queries.beir-nl-webis-touche2020
Dataset Card for BEIR-NL Benchmark
Dataset Summary
BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB).
BEIR-NL contains the following tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-webis-touche2020.webis-touche2020-v3-fatouche2020-fa
Dataset Summary
Touche2020-Fa is a Persian (Farsi) dataset designed for the Retrieval task, specifically focusing on argument retrieval. It is a translated version of the English dataset from the Touché 2020 shared task, included in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection.
Language(s): Persian (Farsi)
Task(s): Retrieval (Argument Retrieval)
Source: Translated from the English Touché 2020 dataset using… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/touche2020-fa.rus-touchewebis-touche2020-top-20-gen-queries
NFCorpus: 20 generated queries (BEIR Benchmark)
This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset.
DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1
id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl).
Questions generated: 20
Code used for generation: evaluate_anserini_docT5query_parallel.py
Below contains the old dataset card for the BEIR benchmark.
Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/webis-touche2020-top-20-gen-queries.gpl-webis-touche2020webis-touche2020-hard-negatives
Dataset Card
Dataset Details
This dataset contains a set of candidate documents for second-stage re-ranking on webis-touche2020
(test split in BEIR). Those candidate documents are composed of hard negatives mined from
gtr-t5-xl as Stage 1 ranker
and ground-truth documents that are known to be relevant to the query. This is a release from our paper
Policy-Gradient Training of Language Models for Ranking, so
please cite it if using this dataset.
Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/NeuralPGRank/webis-touche2020-hard-negatives.splade-webis-touche2020-retrievals
