datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dbpedia-entity
Dataset Card for BEIR Benchmark
dbpedia-entity is one of the datasets from the Entity Retrieval task within BEIR, measuring the retrieval of DbPedia articles for a given query entity.
Dataset Summary
BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/dbpedia-entity.dbpedia-entity-decontaminated
dbpedia-entity (Decontaminated)
A decontaminated version of the dbpedia-entity dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed.
Decontamination methodology
Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files):
Pass 1: Exact hash matching
All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/dbpedia-entity-decontaminated.dbpedia-entity-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/dbpedia-entity-qrels.dbpedia-entity-generated-queries
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/dbpedia-entity-generated-queries.dbpedia-entity.retromae.flex
dbpedia-entity.retromae.flex
Description
RetroMAE index for DBPedia-Entity
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/dbpedia-entity.retromae.flex')
artifact.np_retriever()
Benchmarks
dbpedia-entity/dev
name
nDCG@10
R@1000
np (flat)
0.4687
0.7372
dbpedia-entity/test
name
nDCG@10
R@1000
np (flat)
0.3729
0.678
Reproduction
import pyterrier as pt
from… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/dbpedia-entity.retromae.flex.BEIR-dbpedia-entity-interpretdbpedia-entity.splade-v3.cache
dbpedia-entity.splade-v3.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/dbpedia-entity.splade-v3.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "indexer_cache",
"format": "lz4pickle"… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/dbpedia-entity.splade-v3.cache.dbpedia-entity.terrier
dbpedia-entity.terrier
Description
Terrier index for DBPedia-Entity
Usage
# Load the artifact
import pyterrier as pt
index = pt.Artifact.from_hf('pyterrier/dbpedia-entity.terrier')
index.bm25()
Benchmarks
dbpedia-entity/dev
name
nDCG@10
R@1000
bm25
0.3627
0.7487
dph
0.3617
0.7461
dbpedia-entity/test
name
nDCG@10
R@1000
bm25
0.3039
0.6555
dph
0.305
0.6518
Reproduction
import pyterrier as pt
from tqdm… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/dbpedia-entity.terrier.dbpedia-entity__openai_ada2dbpedia-entity-est
BibTeX entry and citation info
@misc{dorkin2024glilemleveragingglinercontextualized,
title={GliLem: Leveraging GliNER for Contextualized Lemmatization in Estonian},
author={Aleksei Dorkin and Kairit Sirts},
year={2024},
eprint={2412.20597},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.20597},
}
beir-nl-dbpedia-entity
Dataset Card for BEIR-NL Benchmark
Dataset Summary
BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB).
BEIR-NL contains the following tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-dbpedia-entity.dbpedia-entityThis is a reupload of DBpedia-Entity V2 that combines the queries, the relevance judgements, and the corpus (DBpedia dump) in a single place. Please, cite the original authors if you use it.
BibTeX entry and citation info
@inproceedings{Hasibi:2017:DVT,
author = {Hasibi, Faegheh and Nikolaev, Fedor and Xiong, Chenyan and Balog, Krisztian and Bratsberg, Svein Erik and Kotov, Alexander and Callan, Jamie},
title = {DBpedia-Entity V2: A Test Collection for Entity Search}… See the full description on the dataset page: https://huggingface.co/datasets/adorkin/dbpedia-entity.dbpedia-entity
Data Description
Homepage: https://github.com/KID-22/Cocktail
Repository: https://github.com/KID-22/Cocktail
Paper: [Needs More Information]
Dataset Summary
All the 16 benchmarked datasets in Cocktail are listed in the following table.
Dataset
Raw Website
Cocktail Website
Cocktail-Name
md5 for Processed Data
Domain
Relevancy
# Test Query
# Corpus
MS MARCO
Homepage
Homepage
msmarco
985926f3e906fadf0dc6249f23ed850f
Misc.
Binary
6,979
542,203
DL19
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/dbpedia-entity.dbpedia-entity.pisa
dbpedia-entity.pisa
Description
A PISA index for the DBPedia-Entity Dataset
Usage
# Load the artifact
import pyterrier as pt
index = pt.Artifact.from_hf('pyterrier/dbpedia-entity.pisa')
index.bm25() # returns a BM25 retriever
Benchmarks
dbpedia-entity/dev
name
nDCG@10
R@1000
bm25
0.3916
0.7824
dph
0.3961
0.7735
dbpedia-entity/test
name
nDCG@10
R@1000
bm25
0.3271
0.6861
dph
0.3237
0.6794
Reproduction… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/dbpedia-entity.pisa.dbpedia-entity-hard-negatives
Dataset Card
Dataset Details
This dataset contains a set of candidate documents for second-stage re-ranking on dbpedia
(test split in BEIR). Those candidate documents are composed of hard negatives mined from
gtr-t5-xl as Stage 1 ranker
and ground-truth documents that are known to be relevant to the query. This is a release from our paper
Policy-Gradient Training of Language Models for Ranking, so
please cite it if using this dataset.
Direct Use
You can… See the full description on the dataset page: https://huggingface.co/datasets/NeuralPGRank/dbpedia-entity-hard-negatives.beir_dbpedia_entity_test
beir_dbpedia_entity_test
BEIR DBPedia-Entity test split
Field
Value
Benchmark
beir
Sub-benchmark
dbpedia_entity
Type
retrieval
Items
400
Exported from Langfuse.
beir_dbpedia_entity
BEIR DBPedia-Entity (orgrctera/beir_dbpedia_entity)
Overview
DBpedia-Entity v2 is a standard test collection for entity-oriented search over the DBpedia knowledge base: given a short information need expressed in natural language, systems must retrieve DBpedia entities (articles) that satisfy that need. The collection unifies queries from several benchmarks (e.g. SemSearch, INEX, QALD entity search tasks) with graded relevance judgments collected under consistent… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/beir_dbpedia_entity.dbpedia-entity-top-20-gen-queries
NFCorpus: 20 generated queries (BEIR Benchmark)
This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset.
DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1
id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl).
Questions generated: 20
Code used for generation: evaluate_anserini_docT5query_parallel.py
Below contains the old dataset card for the BEIR benchmark.
Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/dbpedia-entity-top-20-gen-queries.dbpedia-entity.splade-v1.cache
dbpedia-entity.splade-v1.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/dbpedia-entity.splade-v1.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "indexer_cache",
"format": "lz4pickle"… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/dbpedia-entity.splade-v1.cache.dbpedia-entity-subgpl-dbpedia-entitybeir_dbpedia-entity
Dataset Card for beir/dbpedia-entity
The beir/dbpedia-entity dataset, provided by the ir-datasets package.
For more information about the dataset, see the documentation.
Data
This dataset provides:
docs (documents, i.e., the corpus); count=4,635,922
queries (i.e., topics); count=467
This dataset is used by: beir_dbpedia-entity_dev, beir_dbpedia-entity_test
Usage
from datasets import load_dataset
docs = load_dataset('irds/beir_dbpedia-entity', 'docs')
for… See the full description on the dataset page: https://huggingface.co/datasets/irds/beir_dbpedia-entity.beir_dbpedia_entity_dev
beir_dbpedia_entity_dev
BEIR DBPedia-Entity dev split
Field
Value
Benchmark
beir
Sub-benchmark
dbpedia_entity
Type
retrieval
Items
67
Exported from Langfuse.
beir_dbpedia-entity_dev
Dataset Card for beir/dbpedia-entity/dev
The beir/dbpedia-entity/dev dataset, provided by the ir-datasets package.
For more information about the dataset, see the documentation.
Data
This dataset provides:
queries (i.e., topics); count=67
qrels: (relevance assessments); count=5,673
For docs, use irds/beir_dbpedia-entity
Usage
from datasets import load_dataset
queries = load_dataset('irds/beir_dbpedia-entity_dev', 'queries')
for record in queries:… See the full description on the dataset page: https://huggingface.co/datasets/irds/beir_dbpedia-entity_dev.beir_dbpedia-entity_test
Dataset Card for beir/dbpedia-entity/test
The beir/dbpedia-entity/test dataset, provided by the ir-datasets package.
For more information about the dataset, see the documentation.
Data
This dataset provides:
queries (i.e., topics); count=400
qrels: (relevance assessments); count=43,515
For docs, use irds/beir_dbpedia-entity
Usage
from datasets import load_dataset
queries = load_dataset('irds/beir_dbpedia-entity_test', 'queries')
for record in queries:… See the full description on the dataset page: https://huggingface.co/datasets/irds/beir_dbpedia-entity_test.
