CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /OpenMathInstruct-1 OpenMathInstruct-1 OpenMathInstruct-1 is a math instruction tuning dataset with 1.8M problem-solution pairs generated using permissively licensed Mixtral-8x7B model. The problems are from GSM8K and MATH training subsets and the solutions are synthetically generated by allowing Mixtral model to use a mix of text reasoning and code blocks executed by Python interpreter. The dataset is split into train and validation subsets that we used in the ablations experiments. These two subsets… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMathInstruct-1.textquestion-answering1M<n<10M254 likes9.8k downloads3y agoHugging Face02nvidia /Nemotron-SpecializedDomains-Finance-v1 Dataset Description Nemotron-SpecializedDomains-Finance is a large-scale synthetic financial question-answering dataset designed to improve LLM performance on specialized financial reasoning and document comprehension tasks. The dataset comprises 326K+ high-quality Q&A pairs generated from SEC filings of S&P 500 companies spanning 2019-2024. This dataset is ready for commercial use. Overview The dataset leverages template-based Synthetic Data Generation (SDG) to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1.texttext-generation100K<n<1M16 likes1.6k downloads7mo agoHugging Face03nvidia /Nemotron-RL-ARC-AGI-v1 Dataset Description: Nemotron-RL-ARC-AGI-v1 is a reinforcement-learning (RL) gym environment dataset of single-step ARC-AGI puzzle prompts intended for RL post-training of large language models. Each row corresponds to one ARC puzzle (a set of (input grid, output grid) demonstration pairs plus a single test input grid) rendered as a text prompt; reward is binary (1.0 / 0.0) determined by exact-match comparison against the ground-truth output grid. No LLM judge is used, no… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-ARC-AGI-v1.texttext-generation10K<n<100K11 likes593 downloads4mo agoHugging Face04nvidia /Nemotron-RL-litmus-bench-v0.1 Dataset Description: Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training. This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-litmus-bench-v0.1.textreinforcement-learning1K<n<10K5 likes385 downloads3mo agoHugging Face05nvidia /AceMath-RewardBenchwebsite | paper AceMath-RewardBench Evaluation Dataset Card The AceMath-RewardBench evaluation dataset evaluates capabilities of a math reward model using the best-of-N (N=8) setting for 7 datasets: GSM8K: 1319 questions Math500: 500 questions Minerva Math: 272 questions Gaokao 2023 en: 385 questions OlympiadBench: 675 questions College Math: 2818 questions MMLU STEM: 3018 questions Each example in the dataset contains: A mathematical question 64 solution attempts with varying… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/AceMath-RewardBench.textquestion-answering10K<n<100K8 likes383 downloads2y agoHugging Face06nvidia /Nemotron-RL-QA-Abstention-v1Nemotron-RL-QA-Abstention-v1 License: cc-by-4.0 Language: en Task Categories: reinforcement-learning, question-answering, text-generation Tags: abstention, question-answering, hotpotqa, software-engineering, health, law, rl, rlvr Configs: default train split at data/train.jsonl Domain: multi-domain question answering, abstention Modality: text Capability Breakdown: Abstention-aware factoid question answering [100%] Source: Hybrid: Automated, Manually Collected, Synthetic Size Bin: <10K… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-QA-Abstention-v1.textreinforcement-learning1K<n<10K5 likes199 downloads3mo agoHugging Face07nvidia /OpenMath-GSM8K-masked OpenMath GSM8K Masked We release a masked version of the GSM8K solutions. This data can be used to aid synthetic generation of additional solutions for GSM8K dataset as it is much less likely to lead to inconsistent reasoning compared to using the original solutions directly. This dataset was used to construct OpenMathInstruct-1: a math instruction tuning dataset with 1.8M problem-solution pairs generated using permissively licensed Mixtral-8x7B model. For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-GSM8K-masked.textquestion-answering1K<n<10K12 likes198 downloads3y agoHugging Face08nvidia /OpenMath-MATH-masked OpenMath GSM8K Masked We release a masked version of the MATH solutions. This data can be used to aid synthetic generation of additional solutions for MATH dataset as it is much less likely to lead to inconsistent reasoning compared to using the original solutions directly. This dataset was used to construct OpenMathInstruct-1: a math instruction tuning dataset with 1.8M problem-solution pairs generated using permissively licensed Mixtral-8x7B model. For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-MATH-masked.textquestion-answering1K<n<10K9 likes150 downloads3y agoHugging Face09agentlans /nvidia-Nemotron-Science-Math NVIDIA Nemotron Science and Math Reasoning This is an unofficial, curated collection derived from NVIDIA's open-source Nemotron datasets. It is designed specifically to train language models in complex scientific and mathematical reasoning by providing structured chain-of-thought (CoT) examples. To ensure efficiency, the shortest available CoT sequence was chosen for each question, filtering out redundant variations while preserving the core logical progression toward the final… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/nvidia-Nemotron-Science-Math.texttext-generation100K<n<1M0 likes59 downloads5mo agoHugging Face105CD-AI /Vietnamese-nvidia-OpenMathInstruct-1-50k-gg-translatedtexttext-generation10K<n<100K7 likes38 downloads3y agoHugging Face11garystafford /fine-tune-nvidia-blackwelltextquestion-answeringn<1K1 likes33 downloads1y agoHugging Face12LLMTeamAkiyama /cleaned_nvidia_OpenCodeReasoning元データ: https://huggingface.co/datasets/nvidia/OpenCodeReasoning データ件数: 11,275 平均トークン数: 11251 最大トークン数: 19,802 合計トークン数: 126,859,041 ファイル形式: JSONL ファイルサイズ: 707.4 MB 難易度スコアが15, カテゴリがcompetition、ライセンスがmitとcc-by-4.0をピックアップ 繰り返し除去 極端に少ない・多いなどを除去 詳しいコードはGithub https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/opencodereasoning tabularquestion-answering10K<n<100K0 likes14 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.