CoolFace
Datasetpublic

irds/antique

Dataset Card for antique The antique dataset, provided by the ir-datasets package. For more information about the dataset, see the documentation. Data This dataset provides: docs (documents, i.e., the corpus); count=403,666 This dataset is used by: antique_test, antique_test_non-offensive, antique_train, antique_train_split200-train, antique_train_split200-valid Usage from datasets import load_dataset docs = load_dataset('irds/antique', 'docs')… See the full description on the dataset page: https://huggingface.co/datasets/irds/antique.

sourceHugging Faceupdated 4y agoView on Hugging Face
0likes25downloads
Dataset Card

Dataset Card for antique

The antique dataset, provided by the ir-datasets package. For more information about the dataset, see the documentation.

Data

This dataset provides:

  • —docs (documents, i.e., the corpus); count=403,666

This dataset is used by: `antique_test`, `antique_test_non-offensive`, `antique_train`, `antique_train_split200-train`, `antique_train_split200-valid`

Usage

python
from datasets import load_dataset

docs = load_dataset('irds/antique', 'docs')
for record in docs:
    record # {'doc_id': ..., 'text': ...}

Note that calling load_dataset will download the dataset (or provide access instructions when it's not public) and make a copy of the data in 🤗 Dataset format.

Citation Information

@inproceedings{Hashemi2020Antique,
  title={ANTIQUE: A Non-Factoid Question Answering Benchmark},
  author={Helia Hashemi and Mohammad Aliannejadi and Hamed Zamani and Bruce Croft},
  booktitle={ECIR},
  year={2020}
}