CoolFace
Datasetpublic

TTimur/boolq_kg

BoolQ (Kyrgyz) This dataset is the Kyrgyz-translated version of the BoolQ benchmark, a reading comprehension task requiring a yes/no answer. ๐Ÿ”๏ธ Part of the KyrgyzLLM-Bench This dataset is a component of the KyrgyzLLM-Bench, a comprehensive suite for evaluating LLMs in Kyrgyz. Main Paper: Bridging the Gap in Less-Resourced Languages: Building a Benchmark for Kyrgyz Language Models Hugging Face Hub: https://huggingface.co/TTimur GitHub Project:โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/TTimur/boolq_kg.

sourceHugging Faceupdated 11mo agoView on Hugging Face
0likes58downloads
Dataset Card

BoolQ (Kyrgyz)

This dataset is the Kyrgyz-translated version of the BoolQ benchmark, a reading comprehension task requiring a yes/no answer.

๐Ÿ”๏ธ Part of the KyrgyzLLM-Bench

This dataset is a component of the KyrgyzLLM-Bench, a comprehensive suite for evaluating LLMs in Kyrgyz.

๐Ÿ“‹ Dataset Description

This benchmark evaluates a model's ability to answer a question about a text passage with a simple "yes" or "no."

  • โ€”Original Dataset: BoolQ
  • โ€”Translation: The dataset was translated using a dual-model machine translation pipeline, followed by expert manual post-editing and quality assurance checks to ensure cultural and linguistic accuracy.

๐Ÿ“œ Citation

If you find this dataset useful in your research, please cite the main project paper:

bibtex
@article{KyrgyzLLM-Bench,
  title={Bridging the Gap in Less-Resourced Languages: Building a Benchmark for Kyrgyz Language Models},
  author={Timur Turatali, Aida Turdubaeva, Islam Zhenishbekov, Zhoomart Suranbaev, Anton Alekseev, Rustem Izmailov},
  year={2025},
  url={https://huggingface.co/datasets/TTimur/boolq_kg}
}