CoolFace
Datasetpublic

MCINext/trec-covid-fa

Dataset Summary TRECCOVID-Fa is a Persian (Farsi) dataset designed for the Retrieval task, specifically focusing on ad-hoc search for COVID-19-related scientific information. It is a translated version of the English dataset from the TREC-COVID shared task, included in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection. Language(s): Persian (Farsi) Task(s): Retrieval (Ad-hoc Search, COVID-19 Information… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/trec-covid-fa.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes19downloads
Dataset Card

Dataset Summary

TRECCOVID-Fa is a Persian (Farsi) dataset designed for the Retrieval task, specifically focusing on ad-hoc search for COVID-19-related scientific information. It is a translated version of the English dataset from the TREC-COVID shared task, included in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection.

  • —Language(s): Persian (Farsi)
  • —Task(s): Retrieval (Ad-hoc Search, COVID-19 Information Retrieval)
  • —Source: Translated from the English TREC-COVID dataset using Google Translate
  • —Part of FaMTEB: Yes — part of the BEIR-Fa collection

Supported Tasks and Leaderboards

This dataset is used to evaluate models' ability to retrieve relevant scientific literature in response to COVID-19 related queries. Performance can be evaluated on the Persian MTEB Leaderboard (filter by language: Persian).

Construction

  • —Translated from the TREC-COVID shared task dataset using the Google Translate API
  • —Based on the CORD-19 corpus (COVID-19 Open Research Dataset)

Translation quality was validated by the FaMTEB team using:

  • —BM25 comparison between original and translated versions
  • —LLM-based evaluation (GEMBA-DA framework) for translation quality

Data Splits

As defined in the FaMTEB paper (Table 5):

  • —Train: 0 samples
  • —Dev: 0 samples
  • —Test: 196,005 samples
Approximate total dataset size: 238k examples (user-provided figure)