CoolFace
Datasetpublic

MCINext/cqadupstack-android-fa

Dataset Summary CQADupstack-android-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) forums. This dataset is a translated version of the "android" (Android Enthusiasts) StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-android-fa.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes32downloads
Dataset Card

Dataset Summary

CQADupstack-android-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) forums. This dataset is a translated version of the "android" (Android Enthusiasts) StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite.

  • Language(s): Persian (Farsi)
  • Task(s): Retrieval (Duplicate Question Retrieval)
  • Source: Translated from English using Google Translate
  • Part of FaMTEB: Yes — under BEIR-Fa

Supported Tasks and Leaderboards

This dataset evaluates text embedding models on their ability to retrieve semantically equivalent or duplicate questions in the Android usage and development domain. Results can be viewed on the Persian MTEB Leaderboard (filter by language: Persian).

Construction

The dataset was created by:

  • Translating the "android" (Android Enthusiasts) subforum from the English CQADupstack dataset using the Google Translate API
  • Preserving the CQA structure for evaluating duplicate retrieval tasks

According to the FaMTEB paper, the BEIR-Fa datasets underwent translation quality evaluation using:

  • BM25 retrieval performance comparisons (English vs Persian)
  • GEMBA-DA evaluation, leveraging LLMs for quality control

Data Splits

As per the FaMTEB paper (Table 5), all CQADupstack-Fa sub-datasets share a pooled evaluation set:

  • Train: 0 samples
  • Dev: 0 samples
  • Test: 480,902 samples (aggregate)

The cqadupstack-android-fa subset contains approximately 25.4k examples (user-provided figure). For more granular splits, consult the dataset provider or FaMTEB documentation.