MCINext/trec-covid-fa
Dataset Summary TRECCOVID-Fa is a Persian (Farsi) dataset designed for the Retrieval task, specifically focusing on ad-hoc search for COVID-19-related scientific information. It is a translated version of the English dataset from the TREC-COVID shared task, included in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection. Language(s): Persian (Farsi) Task(s): Retrieval (Ad-hoc Search, COVID-19 Information… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/trec-covid-fa.
Dataset Summary
TRECCOVID-Fa is a Persian (Farsi) dataset designed for the Retrieval task, specifically focusing on ad-hoc search for COVID-19-related scientific information. It is a translated version of the English dataset from the TREC-COVID shared task, included in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection.
- Language(s): Persian (Farsi)
- Task(s): Retrieval (Ad-hoc Search, COVID-19 Information Retrieval)
- Source: Translated from the English TREC-COVID dataset using Google Translate
- Part of FaMTEB: Yes — part of the BEIR-Fa collection
Supported Tasks and Leaderboards
This dataset is used to evaluate models' ability to retrieve relevant scientific literature in response to COVID-19 related queries. Performance can be evaluated on the Persian MTEB Leaderboard (filter by language: Persian).
Construction
- Translated from the TREC-COVID shared task dataset using the Google Translate API
- Based on the CORD-19 corpus (COVID-19 Open Research Dataset)
Translation quality was validated by the FaMTEB team using:
- BM25 comparison between original and translated versions
- LLM-based evaluation (GEMBA-DA framework) for translation quality
Data Splits
As defined in the FaMTEB paper (Table 5):
- Train: 0 samples
- Dev: 0 samples
- Test: 196,005 samples
Approximate total dataset size: 238k examples (user-provided figure)
