CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-SpecializedDomains-Finance-v1 Dataset Description Nemotron-SpecializedDomains-Finance is a large-scale synthetic financial question-answering dataset designed to improve LLM performance on specialized financial reasoning and document comprehension tasks. The dataset comprises 326K+ high-quality Q&A pairs generated from SEC filings of S&P 500 companies spanning 2019-2024. This dataset is ready for commercial use. Overview The dataset leverages template-based Synthetic Data Generation (SDG) to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1.texttext-generation100K<n<1M16 likes1.6k downloads7mo agoHugging Face02nvidia /Nemotron-RL-ARC-AGI-v1 Dataset Description: Nemotron-RL-ARC-AGI-v1 is a reinforcement-learning (RL) gym environment dataset of single-step ARC-AGI puzzle prompts intended for RL post-training of large language models. Each row corresponds to one ARC puzzle (a set of (input grid, output grid) demonstration pairs plus a single test input grid) rendered as a text prompt; reward is binary (1.0 / 0.0) determined by exact-match comparison against the ground-truth output grid. No LLM judge is used, no… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-ARC-AGI-v1.texttext-generation10K<n<100K11 likes593 downloads4mo agoHugging Face03nvidia /Nemotron-RL-litmus-bench-v0.1 Dataset Description: Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training. This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-litmus-bench-v0.1.textreinforcement-learning1K<n<10K5 likes385 downloads3mo agoHugging Face04nvidia /Nemotron-RL-QA-Abstention-v1Nemotron-RL-QA-Abstention-v1 License: cc-by-4.0 Language: en Task Categories: reinforcement-learning, question-answering, text-generation Tags: abstention, question-answering, hotpotqa, software-engineering, health, law, rl, rlvr Configs: default train split at data/train.jsonl Domain: multi-domain question answering, abstention Modality: text Capability Breakdown: Abstention-aware factoid question answering [100%] Source: Hybrid: Automated, Manually Collected, Synthetic Size Bin: <10K… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-QA-Abstention-v1.textreinforcement-learning1K<n<10K5 likes199 downloads3mo agoHugging Face05jet-ai /ruler-100-nemotron RULER-100 — Nemotron-Nano-v3 tokenized RULER long-context evaluation data, regenerated with the nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (instruct) tokenizer so the labeled context lengths are exact for that model — instead of drifting, as they do when RULER data tokenized for a different model (e.g. Qwen3) is fed to Nemotron. What's here 7 context lengths: 4096, 8192, 16384, 32768, 65536, 131072, 262144 (the model's max). 13 RULER tasks: niah_single_1/2/3… See the full description on the dataset page: https://huggingface.co/datasets/jet-ai/ruler-100-nemotron.tabularquestion-answering10K<n<100K0 likes73 downloads2mo agoHugging Face06amalia-llm /persona_nemotron Persona Nemotron PT Datasets This is a collection of Portuguese synthetic datasets, consisting of 3 datasets, one with general questions from varied topics, one with math questions, and one with instruction-following requests. The prompts were generated using an approach similar to PersonaHub, with a translated version of Nemotron Personas. Both prompts and answers were generated using Gemma 3-27B. This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_nemotron.textquestion-answering100K<n<1M0 likes68 downloads3mo agoHugging Face07agentlans /nvidia-Nemotron-Science-Math NVIDIA Nemotron Science and Math Reasoning This is an unofficial, curated collection derived from NVIDIA's open-source Nemotron datasets. It is designed specifically to train language models in complex scientific and mathematical reasoning by providing structured chain-of-thought (CoT) examples. To ensure efficiency, the shortest available CoT sequence was chosen for each question, filtering out redundant variations while preserving the core logical progression toward the final… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/nvidia-Nemotron-Science-Math.texttext-generation100K<n<1M0 likes59 downloads5mo agoHugging Face08Arsh9210 /Nemotron-RL-litmus-bench-v0.1 Dataset Description: Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training. This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-litmus-bench-v0.1.textreinforcement-learning1K<n<10K0 likes40 downloads2mo agoHugging Face09SkyWhal3 /STXBP1-RAG-Nemotron 🧬⚡ STXBP1-ARIA RAG Database v10.1 - NVIDIA Nemotron Embeddings The most advanced RAG database for STXBP1 therapeutic research. A pre-built ChromaDB vector database containing: 571,816 indexed text chunks from ~17,000 curated PubMed Central (PMC) biomedical papers + 165 base editing analysis entries, (https://huggingface.co/datasets/SkyWhal3/stxbp1-base-editing-sweep), embedded with NVIDIA's state-of-the-art Llama-Nemotron-Embed-1B-v2 model featuring 2048-dimensional embeddings. ⚡… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/STXBP1-RAG-Nemotron.tabulartext-retrievaln<1K0 likes30 downloads8mo agoHugging Face10SkyWhal3 /SNAP25-RAG-Nemotron 🧬⚡ SNAP25-ARIA RAG Database v2.0 - NVIDIA Nemotron Embeddings The first comprehensive RAG database for SNAP25 therapeutic research. A pre-built ChromaDB vector database containing: 76,592 indexed text chunks from 2,043 curated PubMed Central (PMC) biomedical papers + 683 OpenFold3 structural analysis reports + 53 expert-curated knowledge entries, embedded with NVIDIA's state-of-the-art Llama-Nemotron-Embed-1B-v2 model featuring 2048-dimensional embeddings. 🤝 Built for the SNAP25… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/SNAP25-RAG-Nemotron.texttext-retrievaln<1K0 likes28 downloads8mo agoHugging Face11amalia-llm /amalia-Nemotron-SpecializedDomains-Finance-v1 AMALIA Nemotron-SpecializedDomains-Finance-v1 Version of the nvidia/Nemotron-SpecializedDomains-Finance-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to: Remove entries that reference other LLMs or research labs; Remove the reasoning_content field; Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1 This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SpecializedDomains-Finance-v1.texttext-generation100K<n<1M0 likes27 downloads3mo agoHugging Face12Iambackup /Nemotron-RL-litmus-bench-v0.1 Dataset Description: Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training. This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/Iambackup/Nemotron-RL-litmus-bench-v0.1.textreinforcement-learning1K<n<10K0 likes21 downloads3mo agoHugging Face13wflying /nemotron-nano-rl-mcqa-19k Nemotron Nano RL MCQA 19K Nemotron Nano RL MCQA 19K is a 19,670-example English multiple-choice question answering dataset prepared for reinforcement learning with verifiable rewards (RLVR). Its nano_v3_sft_profiled_stem_mcqa identifier and schema correspond to the knowledge-MCQA component of NVIDIA's Nemotron-3-Nano-RL-Training-Blend, represented here in a compact prompt / label / metadata JSONL format. Each record contains a formatted user prompt, the correct option identifier… See the full description on the dataset page: https://huggingface.co/datasets/wflying/nemotron-nano-rl-mcqa-19k.textquestion-answering10K<n<100K0 likes19 downloads2mo agoHugging Face14weblab-GENIAC /aya-ja-nemotron-dpo-maskedgated aya-ja-nemotron-dpo-masked LLMの推論能力を向上させるためのデータセット CohereForAI/aya_datasetから日本語パートを抜粋 deepinfraのnvidia/Nemotron-4-340B-Instructで応答を再生成 2024年8月現在ではnvidia/Nemotron-4-340B-Instructは使用不可 5,651件(6,259件の内608件削除) ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施 Format { "idx": インデックス, "prompt": 日本語の指示文, "chosen": chosenの応答文, "rejected": rejectedの応答文, "chosen_model": chosenとしたモデル, "rejected_model":… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/aya-ja-nemotron-dpo-masked.texttext-generation1K<n<10K4 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.