CoolFace
Datasetpublic

MCINext/cqadupstack-tex-fa

Dataset Summary CQADupstack-tex-Fa is a Persian (Farsi) dataset curated for the Retrieval task, specifically targeting duplicate question detection in community question-answering (CQA) forums. This dataset is a translation of the "TeX - LaTeX" StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source: Translated from… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-tex-fa.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes28downloads
Dataset Card

Dataset Summary

CQADupstack-tex-Fa is a Persian (Farsi) dataset curated for the Retrieval task, specifically targeting duplicate question detection in community question-answering (CQA) forums. This dataset is a translation of the "TeX - LaTeX" StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite.

  • —Language(s): Persian (Farsi)
  • —Task(s): Retrieval (Duplicate Question Retrieval)
  • —Source: Translated from English using Google Translate
  • —Part of FaMTEB: Yes — under BEIR-Fa

Supported Tasks and Leaderboards

This dataset assesses the ability of text embedding models to identify semantically equivalent or duplicate questions within the domain of TeX and LaTeX programming. Evaluation results appear on the Persian MTEB Leaderboard (filter by language: Persian).

Construction

The construction process included:

  • —Translating the "TeX - LaTeX" subforum from the CQADupstack dataset using the Google Translate API
  • —Retaining the original duplicate-retrieval task structure

The FaMTEB paper confirms quality through:

  • —BM25 retrieval comparisons across English and Persian versions
  • —GEMBA-DA framework, which employed LLMs to verify translation fidelity

Data Splits

Per FaMTEB paper (Table 5), all CQADupstack-Fa sub-datasets share the same test pool:

  • —Train: 0 samples
  • —Dev: 0 samples
  • —Test: 480,902 samples (aggregate)

The individual split count for cqadupstack-tex-fa is not provided separately. For detailed partitioning, refer to FaMTEB resources.

Total (user-reported): ~76.2k examples