TTimur/boolq_kg
BoolQ (Kyrgyz) This dataset is the Kyrgyz-translated version of the BoolQ benchmark, a reading comprehension task requiring a yes/no answer. ๐๏ธ Part of the KyrgyzLLM-Bench This dataset is a component of the KyrgyzLLM-Bench, a comprehensive suite for evaluating LLMs in Kyrgyz. Main Paper: Bridging the Gap in Less-Resourced Languages: Building a Benchmark for Kyrgyz Language Models Hugging Face Hub: https://huggingface.co/TTimur GitHub Project:โฆ See the full description on the dataset page: https://huggingface.co/datasets/TTimur/boolq_kg.
BoolQ (Kyrgyz)
This dataset is the Kyrgyz-translated version of the BoolQ benchmark, a reading comprehension task requiring a yes/no answer.
๐๏ธ Part of the KyrgyzLLM-Bench
This dataset is a component of the KyrgyzLLM-Bench, a comprehensive suite for evaluating LLMs in Kyrgyz.
- Main Paper: Bridging the Gap in Less-Resourced Languages: Building a Benchmark for Kyrgyz Language Models
- Hugging Face Hub: https://huggingface.co/TTimur
- GitHub Project: https://github.com/golden-ratio/kyrgyzLLM_bench
๐ Dataset Description
This benchmark evaluates a model's ability to answer a question about a text passage with a simple "yes" or "no."
- Original Dataset: BoolQ
- Translation: The dataset was translated using a dual-model machine translation pipeline, followed by expert manual post-editing and quality assurance checks to ensure cultural and linguistic accuracy.
๐ Citation
If you find this dataset useful in your research, please cite the main project paper:
@article{KyrgyzLLM-Bench,
title={Bridging the Gap in Less-Resourced Languages: Building a Benchmark for Kyrgyz Language Models},
author={Timur Turatali, Aida Turdubaeva, Islam Zhenishbekov, Zhoomart Suranbaev, Anton Alekseev, Rustem Izmailov},
year={2025},
url={https://huggingface.co/datasets/TTimur/boolq_kg}
}