datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mergekit-configs
MergeKit-configs: access all Hub architectures and automate your model merging process
This dataset facilitates the search for compatible architectures for model merging with MergeKit, streamlining the automation of high-performance merge searches. It provides a snapshot of the Hub’s configuration state, eliminating the need to manually open configuration files.
import polars as pl
# Login using e.g. `huggingface-cli login` to access this dataset
df =… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/mergekit-configs.math_mergedTraining dataset contains aime (excluding 2024), math/train, math/test, openai_math_splits/train, and KbsdJames/Omni-MATH/test. Total of 17521 lines of unique problems.
Testing dataset contains aime_24 and math500 (i.e. openai_math_splits/test). Total of 530 lines of unique problems.
math_merged_cot_solA dataset consists problems from flatlander1024/math_merged and cot solutions generated by Llama-3.1-8b-Instruct. The is_correct label indicates whether the solution is correct or not.
Number of lines: 13864, Overall correct rate: 57.3%
merged-cot
MergeCoT Dataset
A large-scale dataset for training models to resolve git merge conflicts using chain-of-thought (CoT) reasoning. This dataset contains 87,690 examples across multiple programming languages, with detailed reasoning traces for merge conflict resolution.
Dataset Summary
MergeCoT provides paired examples of:
Base versions and two conflicting changes (a and b)
Merged results that correctly combine both changes
Chain-of-thought reasoning explaining the merge… See the full description on the dataset page: https://huggingface.co/datasets/chunyoupeng/merged-cot.math_merged_deduped_OR1_dapo
Math subset for training L1 using RL
This dataset is inspired by LLM360/Reasoning360(GURU92K-math), but reproduced from DAPO-Math-17K and Skywork-OR1-Math. DeepScaleR was not used for source duplications.
Dataset Details
Dataset Description
Curated by: Leon (Me)
Funded by [optional]: AIGCode/Koting Intelligence
Language(s) (NLP): Mostly in English with a few in Chinese
License: MIT (following GURU-92K)
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Leon-Leee/math_merged_deduped_OR1_dapo.cybersec-reasoning-merged
Cybersecurity Reasoning Dataset (Merged)
Dataset Description
This dataset combines two high-quality cybersecurity reasoning datasets to create a comprehensive resource for training language models on security-related tasks with chain-of-thought reasoning.
Dataset Summary
Total Samples: 23,146
Languages: English
Format: Instruction-following with explicit reasoning chains
Domain: Cybersecurity (vulnerabilities, CVE/CWE mapping, security analysis)… See the full description on the dataset page: https://huggingface.co/datasets/Mohannadcse/cybersec-reasoning-merged.PyThagoreans-Merged
PyThagoreans Dataset
Overview
The PyThagoreans dataset is a comprehensive collection of math problems and their solutions, designed to assist in learning and practicing mathematical problem-solving. This dataset includes a variety of problems, expected answers, and predicted answers, making it a valuable resource for students, educators, and researchers.
Dataset Details
Modalities
Text: The dataset primarily contains text data, including math… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/PyThagoreans-Merged.Material-mechanics-merge
Material-mechanics-merge
Dataset Overview
Material-mechanics-merge is an expanded Chinese-language instruction dataset for materials mechanics and engineering mechanics. It provides question--answer examples for building and evaluating domain-focused educational language models.
Dataset Details
Maintainer: CYHcyh66
Language: Chinese
License: Apache-2.0
Format: JSON
Split: train
Size: 774 examples
Task: Instruction-following question answering… See the full description on the dataset page: https://huggingface.co/datasets/CYHcyh66/Material-mechanics-merge.math_merged_cot_sol_hardA dataset consists the hard problems (aime + problems with level >= 5) from flatlander1024/math_merged and cot solutions generated by Qwen2.5-32B-Instruct. The is_correct label indicates whether the solution is correct or not.
Number of lines: 6940, Overall correct rate: 45.4%
merged-medical-qatombench_merged
TomBench Merged Dataset (Exact Matching)
This dataset contains the merged results of TomBench evaluation with the original TomBench dataset, using exact string matching.
Dataset Statistics
Total records: 2860
Exact matches: 2860
Manual matches: 0
Average model score: 0.5066
Matching Strategy
This version uses exact string matching after text normalization:
Remove extra whitespace and normalize formatting
Match stories exactly between datasets
Report any… See the full description on the dataset page: https://huggingface.co/datasets/ycfNTU/tombench_merged.Alpaca_orca_bongchat_mergedswesat-skolprov-merged
SweSAT + Swedish Skolprov (Merged Dataset)
Dataset Description
This dataset is a unified, state-of-the-art benchmark designed for evaluating Large Language Models (LLMs) on Swedish text comprehension, vocabulary, and logical reasoning. It is constructed by merging two prominent Swedish test datasets:
SweSAT-1.0: Questions sourced from the Swedish University Entrance Exam (Högskoleprovet) spanning from 2020-10-25 to 2024-04-13.
Swedish Skolprov: Diverse Swedish academic… See the full description on the dataset page: https://huggingface.co/datasets/jonasaise/swesat-skolprov-merged.re-merged-pf-2swesat-skolprov-superlim-merged
Dataset Card for the Swedish NLU Benchmark Collection
Dataset Description
This dataset is a comprehensive, deduplicated benchmark collection specifically designed for evaluating the Swedish Natural Language Understanding (NLU) capabilities of Large Language Models (LLMs). The dataset merges high-quality multiple-choice scholastic examinations with a diverse suite of NLP and reasoning tasks.
The benchmark contains over 450,000 unique queries compiled into a unified JSONL… See the full description on the dataset page: https://huggingface.co/datasets/jonasaise/swesat-skolprov-superlim-merged.merged_math_test
Merged Math Test Dataset
This dataset merges multiple math competition and benchmark datasets into a unified format with three fields:
problem: The math problem statement
answer: The answer to the problem
source: The source dataset
Source Datasets
olympiadbench (674 examples): math-ai/olympiadbench
Mapped: question → problem, final_answer[0] → answer
math500 (500 examples): math-ai/math500
Mapped: problem → problem, answer → answer
aime25 (30 examples):… See the full description on the dataset page: https://huggingface.co/datasets/chaosc/merged_math_test.merged_aime_test
Merged AIME Test Dataset
This dataset merges AIME 2024 and AIME 2025 datasets into a unified format with three fields:
problem: The math problem statement
answer: The answer to the problem
source: The source dataset
Source Datasets
aime25 (30 examples): math-ai/aime25
Mapped: problem → problem, answer → answer
aime24 (30 examples): math-ai/aime24
Mapped: problem → problem, extracted answer from \boxed{} in solution → answer
Total Statistics
Total… See the full description on the dataset page: https://huggingface.co/datasets/chaosc/merged_aime_test.merged-pfmerged_final2
Merged Final 2
This dataset is a merged version of Pamzyy's QA dataset and the Sinhala NSINA news dataset.
Each row contains either:
a QA pair from the original dataset, or
a headline + news content converted into question and answer format.
