datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BenchMAX_Problem_Solving
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Problem_Solving is a dataset of BenchMAX, sourcing from LiveCodeBench_v4, which evaluates the code generation capability for solving multilingual competitive code problems.
We extend the original English dataset by 16 non-English languages.
The… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Problem_Solving.formal_problem_solving_main
Dataset Card for Formal Problem-Solving Benchmarks
This dataset is part of the official implementation of Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving, accepted as an ICML 2026 Spotlight paper.
Links:
Paper: https://openreview.net/forum?id=hgMZraPlSv
Project: https://github.com/Purewhite2019/formal_problem_solving_main
Overview
The benchmark supports three evaluation settings:
Formal Problem-Solving (FPS): Given a… See the full description on the dataset page: https://huggingface.co/datasets/purewhite42/formal_problem_solving_main.Reasoning_Problem_Solving_Dataset
Reasoning and Problem-Solving Dataset (RPSD)
Overview
The Reasoning and Problem-Solving Dataset (RPSD) is a comprehensive, high-quality set of synthetically generated question-answer pairs (150k+) tailored for training AI systems in logical reasoning and problem-solving. It spans multiple domains, including core reasoning techniques, specialized fields like science, mathematics, engineering, computer science, and philosophy, along with practical, real-world… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/Reasoning_Problem_Solving_Dataset.Problem-Solving-Insights-Based-on-Kazakh-Traditions
🇰🇿 Problem-Solving Insights Based on Kazakh Traditions
📖 Overview
Problem-Solving Insights Based on Kazakh Traditions is a instruction-tuning dataset designed to bridge the gap between ancient Kazakh wisdom and modern societal challenges.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
8,005
Total Words (approx.)
4,030,925
Avg. Words per Sample
503
Word Count Distribution (Per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Problem-Solving-Insights-Based-on-Kazakh-Traditions.ToT_Reasoning_Problem_Solving_Dataset_V2
ToT-RPSD-V2
This dataset consists of 70,000 high-quality, synthetically generated Q&A pairs with a strong emphasis on reasoning (inspired by o1 type reasoning) and the use of "Train of Thought" methodologies. Each entry is meticulously structured into six key components: the question, answer, reasoning (detailing the thought process leading to the answer), a unique ID, topic tags, and a difficulty level. While the dataset strongly focuses on science and cognitive tasks, it… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/ToT_Reasoning_Problem_Solving_Dataset_V2.
