CoolFace
Datasetpublic

MCINext/cqadupstack-stats-fa

Dataset Summary CQADupstack-stats-Fa is a Persian (Farsi) dataset developed for the Retrieval task, focusing on duplicate question detection in community question-answering (CQA) platforms. This dataset is a translated version of the "Cross Validated" (Stats) subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source: Translated from… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-stats-fa.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes20downloads
Dataset Card

Dataset Summary

CQADupstack-stats-Fa is a Persian (Farsi) dataset developed for the Retrieval task, focusing on duplicate question detection in community question-answering (CQA) platforms. This dataset is a translated version of the "Cross Validated" (Stats) subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite.

  • —Language(s): Persian (Farsi)
  • —Task(s): Retrieval (Duplicate Question Retrieval)
  • —Source: Translated from English using Google Translate
  • —Part of FaMTEB: Yes — under BEIR-Fa

Supported Tasks and Leaderboards

This dataset evaluates text embedding models on their ability to retrieve semantically similar or duplicate questions in domains such as statistics, machine learning, and data analysis. Results can be viewed on the Persian MTEB Leaderboard (filter by language: Persian).

Construction

The dataset was generated through:

  • —Translating the "stats" StackExchange subforum from CQADupstack via Google Translate API
  • —Retaining the original CQA structure for duplicate retrieval

According to the FaMTEB paper, the BEIR-Fa datasets underwent quality validation via:

  • —BM25 retrieval performance comparisons (English vs Persian)
  • —GEMBA-DA evaluation, leveraging LLMs for translation quality assurance

Data Splits

As per the FaMTEB paper (Table 5), all CQADupstack-Fa sub-datasets are part of a shared evaluation pool:

  • —Train: 0 samples
  • —Dev: 0 samples
  • —Test: 480,902 samples (aggregate)

The exact subset size for cqadupstack-stats-fa is approximately 43.8k examples (user-provided). For detailed splits, consult the FaMTEB dataset repository or maintainers.