CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01louisbrulenaudet /mergekit-configs MergeKit-configs: access all Hub architectures and automate your model merging process This dataset facilitates the search for compatible architectures for model merging with MergeKit, streamlining the automation of high-performance merge searches. It provides a snapshot of the Hub’s configuration state, eliminating the need to manually open configuration files. import polars as pl # Login using e.g. `huggingface-cli login` to access this dataset df =… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/mergekit-configs.tabularquestion-answering100K<n<1M8 likes205 downloads2y agoHugging Face02flatlander1024 /math_mergedTraining dataset contains aime (excluding 2024), math/train, math/test, openai_math_splits/train, and KbsdJames/Omni-MATH/test. Total of 17521 lines of unique problems. Testing dataset contains aime_24 and math500 (i.e. openai_math_splits/test). Total of 530 lines of unique problems. textquestion-answering10K<n<100K0 likes54 downloads1y agoHugging Face03flatlander1024 /math_merged_cot_solA dataset consists problems from flatlander1024/math_merged and cot solutions generated by Llama-3.1-8b-Instruct. The is_correct label indicates whether the solution is correct or not. Number of lines: 13864, Overall correct rate: 57.3% textquestion-answering10K<n<100K0 likes48 downloads1y agoHugging Face04chunyoupeng /merged-cot MergeCoT Dataset A large-scale dataset for training models to resolve git merge conflicts using chain-of-thought (CoT) reasoning. This dataset contains 87,690 examples across multiple programming languages, with detailed reasoning traces for merge conflict resolution. Dataset Summary MergeCoT provides paired examples of: Base versions and two conflicting changes (a and b) Merged results that correctly combine both changes Chain-of-thought reasoning explaining the merge… See the full description on the dataset page: https://huggingface.co/datasets/chunyoupeng/merged-cot.texttext-generation10K<n<100K0 likes43 downloads10mo agoHugging Face05Leon-Leee /math_merged_deduped_OR1_dapo Math subset for training L1 using RL This dataset is inspired by LLM360/Reasoning360(GURU92K-math), but reproduced from DAPO-Math-17K and Skywork-OR1-Math. DeepScaleR was not used for source duplications. Dataset Details Dataset Description Curated by: Leon (Me) Funded by [optional]: AIGCode/Koting Intelligence Language(s) (NLP): Mostly in English with a few in Chinese License: MIT (following GURU-92K) Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Leon-Leee/math_merged_deduped_OR1_dapo.textquestion-answering100K<n<1M0 likes35 downloads1y agoHugging Face06Mohannadcse /cybersec-reasoning-merged Cybersecurity Reasoning Dataset (Merged) Dataset Description This dataset combines two high-quality cybersecurity reasoning datasets to create a comprehensive resource for training language models on security-related tasks with chain-of-thought reasoning. Dataset Summary Total Samples: 23,146 Languages: English Format: Instruction-following with explicit reasoning chains Domain: Cybersecurity (vulnerabilities, CVE/CWE mapping, security analysis)… See the full description on the dataset page: https://huggingface.co/datasets/Mohannadcse/cybersec-reasoning-merged.textquestion-answering10K<n<100K2 likes35 downloads9mo agoHugging Face07prithivMLmods /PyThagoreans-Merged PyThagoreans Dataset Overview The PyThagoreans dataset is a comprehensive collection of math problems and their solutions, designed to assist in learning and practicing mathematical problem-solving. This dataset includes a variety of problems, expected answers, and predicted answers, making it a valuable resource for students, educators, and researchers. Dataset Details Modalities Text: The dataset primarily contains text data, including math… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/PyThagoreans-Merged.textquestion-answering1M<n<10M2 likes32 downloads2y agoHugging Face08CYHcyh66 /Material-mechanics-merge Material-mechanics-merge Dataset Overview Material-mechanics-merge is an expanded Chinese-language instruction dataset for materials mechanics and engineering mechanics. It provides question--answer examples for building and evaluating domain-focused educational language models. Dataset Details Maintainer: CYHcyh66 Language: Chinese License: Apache-2.0 Format: JSON Split: train Size: 774 examples Task: Instruction-following question answering… See the full description on the dataset page: https://huggingface.co/datasets/CYHcyh66/Material-mechanics-merge.textquestion-answeringn<1K0 likes32 downloads26d agoHugging Face09flatlander1024 /math_merged_cot_sol_hardA dataset consists the hard problems (aime + problems with level >= 5) from flatlander1024/math_merged and cot solutions generated by Qwen2.5-32B-Instruct. The is_correct label indicates whether the solution is correct or not. Number of lines: 6940, Overall correct rate: 45.4% textquestion-answering1K<n<10K0 likes27 downloads1y agoHugging Face10Faithality /merged-medical-qatextquestion-answering10K<n<100K0 likes21 downloads1y agoHugging Face11ycfNTU /tombench_merged TomBench Merged Dataset (Exact Matching) This dataset contains the merged results of TomBench evaluation with the original TomBench dataset, using exact string matching. Dataset Statistics Total records: 2860 Exact matches: 2860 Manual matches: 0 Average model score: 0.5066 Matching Strategy This version uses exact string matching after text normalization: Remove extra whitespace and normalize formatting Match stories exactly between datasets Report any… See the full description on the dataset page: https://huggingface.co/datasets/ycfNTU/tombench_merged.tabularquestion-answering1K<n<10K0 likes20 downloads1y agoHugging Face12ahnaf702 /Alpaca_orca_bongchat_mergedtextquestion-answering100K<n<1M0 likes17 downloads2y agoHugging Face13jonasaise /swesat-skolprov-merged SweSAT + Swedish Skolprov (Merged Dataset) Dataset Description This dataset is a unified, state-of-the-art benchmark designed for evaluating Large Language Models (LLMs) on Swedish text comprehension, vocabulary, and logical reasoning. It is constructed by merging two prominent Swedish test datasets: SweSAT-1.0: Questions sourced from the Swedish University Entrance Exam (Högskoleprovet) spanning from 2020-10-25 to 2024-04-13. Swedish Skolprov: Diverse Swedish academic… See the full description on the dataset page: https://huggingface.co/datasets/jonasaise/swesat-skolprov-merged.textmultiple-choicen<1K0 likes16 downloads7mo agoHugging Face14baebee /re-merged-pf-2textquestion-answering10K<n<100K0 likes14 downloads3y agoHugging Face15jonasaise /swesat-skolprov-superlim-merged Dataset Card for the Swedish NLU Benchmark Collection Dataset Description This dataset is a comprehensive, deduplicated benchmark collection specifically designed for evaluating the Swedish Natural Language Understanding (NLU) capabilities of Large Language Models (LLMs). The dataset merges high-quality multiple-choice scholastic examinations with a diverse suite of NLP and reasoning tasks. The benchmark contains over 450,000 unique queries compiled into a unified JSONL… See the full description on the dataset page: https://huggingface.co/datasets/jonasaise/swesat-skolprov-superlim-merged.textquestion-answering100K<n<1M0 likes10 downloads7mo agoHugging Face16chaosc /merged_math_test Merged Math Test Dataset This dataset merges multiple math competition and benchmark datasets into a unified format with three fields: problem: The math problem statement answer: The answer to the problem source: The source dataset Source Datasets olympiadbench (674 examples): math-ai/olympiadbench Mapped: question → problem, final_answer[0] → answer math500 (500 examples): math-ai/math500 Mapped: problem → problem, answer → answer aime25 (30 examples):… See the full description on the dataset page: https://huggingface.co/datasets/chaosc/merged_math_test.textquestion-answering1K<n<10K0 likes7 downloads9mo agoHugging Face17chaosc /merged_aime_test Merged AIME Test Dataset This dataset merges AIME 2024 and AIME 2025 datasets into a unified format with three fields: problem: The math problem statement answer: The answer to the problem source: The source dataset Source Datasets aime25 (30 examples): math-ai/aime25 Mapped: problem → problem, answer → answer aime24 (30 examples): math-ai/aime24 Mapped: problem → problem, extracted answer from \boxed{} in solution → answer Total Statistics Total… See the full description on the dataset page: https://huggingface.co/datasets/chaosc/merged_aime_test.textquestion-answeringn<1K0 likes7 downloads9mo agoHugging Face18baebee /merged-pftextquestion-answering10K<n<100K0 likes5 downloads3y agoHugging Face19Pamzyy /merged_final2gated Merged Final 2 This dataset is a merged version of Pamzyy's QA dataset and the Sinhala NSINA news dataset. Each row contains either: a QA pair from the original dataset, or a headline + news content converted into question and answer format. textquestion-answering1M<n<10M0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.