CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-Post-Training-Dataset-v1 Nemotron-Post-Training-Dataset-v1 Release This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5. Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency. Efficiency… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1.text10M<n<100M194 likes14k downloads1y agoHugging Face02nvidia /Nemotron-Post-Training-Dataset-v2gated Nemotron-Post-Training-Dataset-v2 Release Data Overview This dataset adds to NVIDIA’s post-training dataset releases with an extension of SFT and RL data into five target languages: Spanish, French, German, Italian and Japanese. The data supports improvements of math, code, general reasoning, and instruction following capabilities of the NVIDIA-Nemotron-Nano-9B-v2-Base, in support of release of NVIDIA-Nemotron-Nano-8B-v2-Reasoning. NVIDIA-Nemotron-Nano-9B is a family of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2.text1M<n<10M156 likes5.5k downloads1y agoHugging Face03Post-training-Data-Flywheel /gorilla-openfunctions-v1text10K<n<100K0 likes1.7k downloads2y agoHugging Face04openeurollm /Nemotron-Post-Training-Dataset-v2-decontaminated Decontamination This dataset is a decontaminated version of nvidia/Nemotron-Post-Training-Dataset-v2. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Nemotron-Post-Training-Dataset-v2-decontaminated.text1M<n<10M1 likes1.1k downloads6mo agoHugging Face05Post-training-Data-Flywheel /Salesforce-xlam-function-calling-60ktext10K<n<100K0 likes774 downloads2y agoHugging Face06ghostcc3 /mix-context-post-training-128k Mix-Context Post-Training Dataset for 128K Context Extension Overview Mix-Context Post-Training 128K is a dataset designed specifically for post-training context window extension of pretrained LLMs. It targets the stage after base pretraining, where a model is adapted to operate over much longer contexts (up to 128K tokens) while preserving short-context behavior. The dataset mixes short- and long-context packed sequences with a controlled length distribution to support:… See the full description on the dataset page: https://huggingface.co/datasets/ghostcc3/mix-context-post-training-128k.text-generation10K<n<100K3 likes733 downloads8mo agoHugging Face07MaziyarPanahi /Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT Llama-Nemotron-Post-Training-Dataset-v1 in ShareGPT Format This dataset is a conversion of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1 into the ShareGPT format while preserving the original splits and columns. Format Each example contains all original fields plus a messages array: { "input": "original input text", "output": "original output text", ... (other original columns) ..., "messages": [ {"role": "user", "content": "User message"}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT.text10M<n<100M41 likes702 downloads1y agoHugging Face08Podtech /Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt Dataset Overview This dataset is a reformatted subset of the tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1 dataset, specifically derived from the v1-Ja-202601 subset. It was created to facilitate Continuous Pre-Training (CPT) by extracting only the text_gpt_oss field from the original data. Dataset Statistics & Token Counts The token counts for each category were calculated using the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt.texttext-generation1M<n<10M0 likes527 downloads1mo agoHugging Face09MaziyarPanahi /Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1 converted to ShareGPT format and merged into a single dataset. Format Each example contains all original fields plus a messages array: { "input": "original input text", "output": "original output text", ... (other original columns) ..., "original_split": "code|math|science|chat|safety", "messages": [ {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT.text1M<n<10M3 likes351 downloads2y agoHugging Face10Post-training-Data-Flywheel /AutoIF-instruct-61ktext10K<n<100K18 likes260 downloads2y agoHugging Face11MaziyarPanahi /Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1 converted to ShareGPT format and merged into a single dataset. Format Each example contains all original fields plus a messages array: { "input": "original input text", "output": "original output text", ... (other original columns) ..., "original_split": "code|math|science|chat|safety", "messages": [ {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT.text1M<n<10M3 likes244 downloads2y agoHugging Face12LumiOpen /Llama-Nemotron-Post-Training-Dataset-SFT-math-FI Llama-Nemotron-Post-Training-Dataset-SFT-math-FI This dataset is a Finnish machine-translated version of the SFT/math split from the original nvidia/Llama-Nemotron-Post-Training-Dataset. The data was created by translating the original English math SFT subset into Finnish using the DeepSeek-V3 model. Translation Process The user prompt and the thinking traces were translated separately in two LLM requests. For traces, the <think> and </think> tokens were preserved… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/Llama-Nemotron-Post-Training-Dataset-SFT-math-FI.texttext-generation1M<n<10M1 likes244 downloads2mo agoHugging Face13kurakurai /Luth-2-Post-Training-SFT Luth-2-Post-Training-SFT Luth-2-Post-Training-SFT is the French supervised fine-tuning mixture used to train Luth-2-0.8B and Luth-2-2B. It spans math, code, knowledge, instruction following and tool calling in a single schema, with 1,969,768 examples and 3.12B training tokens. 📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD 🤗 Models: Luth-2-0.8B · Luth-2-2B 📊 Datasets: SFT · RL 💻 Code: GitHub 🏆 Leaderboard: French LLM Leaderboard Composition… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-SFT.texttext-generation1M<n<10M5 likes209 downloads2mo agoHugging Face14HuggingFaceTB /post-training-benchmarks-viewertabularn<1K3 likes187 downloads11mo agoHugging Face15Post-training-Data-Flywheel /AutoIF-instruct-61k-with-funcstext10K<n<100K8 likes185 downloads2y agoHugging Face16typhoon-ai /typhoon-s-instruct-post-training Typhoon-S Instruct Post-Training Dataset Summary This dataset is a post-training corpus used in the Typhoon-S recipe for building Sovereign AI: high-performing, region- and domain-specific LLMs that remain localized, controllable, and resource-efficient. It is designed to help transform a sovereignty-adapted base model into a capable assistant while preserving target-language strengths. The dataset follows a two-part mixture philosophy: Target-language (Thai) alignment… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-instruct-post-training.texttext-generation100K<n<1M0 likes182 downloads8mo agoHugging Face17cmu-lti /osim-post-training SOUL This CMU-LTI mirror hosts the post-training data used for ODYSSIM releases. It mirrors the original sunweiwei/Soul dataset layout under the CMU-LTI organization. SOUL is the data suite for human behavior simulation used in Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO), spanning conversation, social simulation, social cognition, role-play, and human-centric evaluation. 📄 Paper: https://arxiv.org/abs/2605.20506 💻 Code:… See the full description on the dataset page: https://huggingface.co/datasets/cmu-lti/osim-post-training.texttext-generation10K<n<100K1 likes181 downloads4mo agoHugging Face18minzh23 /Qwen3-8B-Nemotron-Post-Training-Dataset-v2text1M<n<10M1 likes157 downloads6mo agoHugging Face19ZeroAgency /mistral-nvidia-Llama-Nemotron-Post-Training-Dataset-sfttext10M<n<100M0 likes143 downloads1y agoHugging Face20AIGym /post-training-v1text100K<n<1M0 likes140 downloads1y agoHugging Face21hamishivi /llama_nemotron_post_training_sft_sciencetext100K<n<1M0 likes136 downloads1y agoHugging Face22ChavyvAkvar /Nemotron-Post-Training-Dataset-v2-chat-Convertedtext100K<n<1M0 likes130 downloads1y agoHugging Face23kurakurai /Luth-2-Post-Training-RL Luth-2-Post-Training-RL Luth-2-Post-Training-RL is the French RL prompt collection used to post-train Luth-2-0.8B and Luth-2-2B. 📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD 🤗 Models: Luth-2-0.8B · Luth-2-2B 📊 Datasets: SFT · RL 💻 Code: GitHub 🏆 Leaderboard: French LLM Leaderboard Composition Config Rows Verifier fields math 20,000 prompt, solution math_hard 17,888 prompt, solution code 46,661 prompt, unit_tests… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-RL.texttext-generation100K<n<1M4 likes128 downloads2mo agoHugging Face24nick007x /Nemotron-Post-Training-Dataset-v1 Nemotron-Post-Training-Dataset-v1 Release This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5. Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency.… See the full description on the dataset page: https://huggingface.co/datasets/nick007x/Nemotron-Post-Training-Dataset-v1.text10M<n<100M0 likes124 downloads6mo agoHugging Face25ankitdhiman /nemotron-post-training-dataset-v1-processedchat and tools subset of nvidia/Nemotron-Post-Training-Dataset-v1 without thinking tokens and user messages filled. text1M<n<10M1 likes117 downloads1y agoHugging Face26Post-training-Data-Flywheel /OpenOrcatext1M<n<10M0 likes114 downloads2y agoHugging Face27saurabh5 /nemotron-post-training-dataset-v1-chattext100K<n<1M1 likes113 downloads1y agoHugging Face28minzh23 /Qwen3-4B-Nemotron-Post-Training-Dataset-v2text1M<n<10M0 likes113 downloads6mo agoHugging Face29Post-training-Data-Flywheel /NousResearch-hermes-function-calling-v1text1K<n<10K0 likes88 downloads2y agoHugging Face30Post-training-Data-Flywheel /gorilla-apibenchtext10K<n<100K0 likes82 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.