CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-SFT-Instruction-Following-Chat-v3 Dataset Description: The Nemotron-Instruction-Following-Chat-v3 dataset is designed to strengthen multi-turn, interactive capabilities, including open-ended chat and precise instruction following. The chat subset uses human written prompts from sources like lmarena, lmsys, and wildchat as seed prompts. Responses are generated with GLM-5. Multiple responses are sampled from the model and the best response as judged by pairwise comparisons using Qwen3-Nemotron-235B-A22B-GenRM-2603… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.texttext-generation100K<n<1M20 likes5.4k downloads4mo agoHugging Face02nvidia /Nemotron-Instruction-Following-Chat-v1 Dataset Description: The Nemotron-Instruction-Following-Chat-v1 dataset is designed to broadly strengthen the model’s interactive capabilities, spanning open-ended chat, precise instruction following, and reliable structured output generation. It combines refreshed chat data from Nemotron-Post-Training-Dataset-v2 (extended to multi-turn) with synthetic dialogues produced by strong frontier models such as GPT-OSS-120B and Qwen3-235B variants. This dataset is ready for commercial… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Instruction-Following-Chat-v1.text100K<n<1M130 likes1.1k downloads9mo agoHugging Face03Chinese-Vicuna /instruct_chat_50k.jsonlinstruct_chat_50k.jsonl which is composed of 30k Chinese sharegpt dataset and 20k alpaca-instruction-Chinese-dataset textquestion-answering10K<n<100K44 likes110 downloads3y agoHugging Face04Chat-Error /Long-instructionstext10K<n<100K3 likes107 downloads3y agoHugging Face05amalia-llm /amalia-Nemotron-Instruction-Following-Chat-v1 AMALIA Nemotron-Instruction-Following-Chat-v1 Version of the nvidia/Nemotron-Instruction-Following-Chat-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to: Remove entries where the capability_target field was 'chat' from the chat_if split; Remove entries that reference other LLMs or research labs; Remove entries that contained the string '/imagine prompt:'; Removed the reasoning_content field; Original… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-Instruction-Following-Chat-v1.text10K<n<100K0 likes64 downloads3mo agoHugging Face06aiuser3993 /Llama-Krikri-8B-Instruct-ChatCreative and vocabulary-rich (LLM-focused) Greek conversation data generated by ilsp/Llama-Krikri-8B-Instruct-GGUF for distlization purposes. textn<1K0 likes63 downloads26d agoHugging Face07Fredithefish /Instruction-Tuning-with-GPT-4-RedPajama-Chat Instruction Tuning with GPT 4 RedPajama-Chat This dataset has been converted from the Instruction-Tuning-with-GPT-4 dataset for the purpose of fine-tuning the RedPajama-INCITE-Chat-3B-v1 model. About Instruction-Tuning-with-GPT-4 English Instruction-Following Data generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. Usage and License Notices The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Instruction-Tuning-with-GPT-4-RedPajama-Chat.textquestion-answering10K<n<100K6 likes58 downloads3y agoHugging Face08Jackrong /Qwen3-235B-A22B-Instruct-2507-Distilled-chat Qwen3-235B-A22B-Instruct-2507-Distilled-chat📚 Curated/Funded/Shared by: [Jack Rong] Language(s): English (major), Chinese, Русский, 한국어, 日本語, others License: [apache-2.0] Distilled Model: 🏆Qwen/Qwen3-235B-A22B-Instruct-2507 Qwen3-235B-A22B-Instruct-2507 Benchmarks📊 Introduction: The objectives of this project are: Focus on chat capabilities (excluding CoT), covering cross-lingual real-world Q&A/explanation/generation; Utilize… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Qwen3-235B-A22B-Instruct-2507-Distilled-chat.texttable-question-answering1K<n<10K6 likes57 downloads1y agoHugging Face09Sidsidney /Nemotron-Instruction-Following-Chat-v1 Dataset Description: The Nemotron-Instruction-Following-Chat-v1 dataset is designed to broadly strengthen the model’s interactive capabilities, spanning open-ended chat, precise instruction following, and reliable structured output generation. It combines refreshed chat data from Nemotron-Post-Training-Dataset-v2 (extended to multi-turn) with synthetic dialogues produced by strong frontier models such as GPT-OSS-120B and Qwen3-235B variants. This dataset is ready for commercial… See the full description on the dataset page: https://huggingface.co/datasets/Sidsidney/Nemotron-Instruction-Following-Chat-v1.text100K<n<1M0 likes37 downloads9mo agoHugging Face10amalia-llm /amalia-Nemotron-SFT-Instruction-Following-Chat-v2 Nemotron-SFT-Instruction-Following-Chat-v2 Version of the nvidia/Nemotron-SFT-Instruction-Following-Chat-v2 dataset used in the AMALIA's Supervised Fine-Tuning stage, both in the base and ramp down stages. The ramp down stage comprised a randomly select subset of the base subset, where part was translated to European Portuguese using google/gemma-4-31B-it. This dataset went through a processing pipeline to: Remove entries that reference other LLMs or research labs;… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Instruction-Following-Chat-v2.texttext-generation10K<n<100K0 likes30 downloads3mo agoHugging Face11open-llm-leaderboard /BenevolenceMessiah__Yi-Coder-9B-Chat-Instruct-TIES-MoE-v1.0-detailsgated Dataset Card for Evaluation run of BenevolenceMessiah/Yi-Coder-9B-Chat-Instruct-TIES-MoE-v1.0 Dataset automatically created during the evaluation run of model BenevolenceMessiah/Yi-Coder-9B-Chat-Instruct-TIES-MoE-v1.0 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BenevolenceMessiah__Yi-Coder-9B-Chat-Instruct-TIES-MoE-v1.0-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face12Wenhao97 /gpt4o-mini-instruction-synthesis-chat-formattext1K<n<10K0 likes10 downloads2y agoHugging Face13Shaer-AI /shaer-eval-instruction-yehia-base-sft-chat-template shaer-eval-instruction-yehia-base-sft-chat-template Clean paper-facing evaluation dataset for the Shaer benchmark. Rows: 3481 Split: test Schema id base_meter form requested_bayts requested_num_lines description enhanced_description reference_completion generated_text meter count_adherence description_adherence meaning fluency coherence poeticness Notes meter is the row-level metrical conformity score used in the paper. This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/shaer-eval-instruction-yehia-base-sft-chat-template.tabular1K<n<10K0 likes4 downloads3mo agoHugging Face14Rendra86318 /Nemotron-Instruction-Following-Chat-v1 Dataset Description: The Nemotron-Instruction-Following-Chat-v1 dataset is designed to broadly strengthen the model’s interactive capabilities, spanning open-ended chat, precise instruction following, and reliable structured output generation. It combines refreshed chat data from Nemotron-Post-Training-Dataset-v2 (extended to multi-turn) with synthetic dialogues produced by strong frontier models such as GPT-OSS-120B and Qwen3-235B variants. This dataset is ready for commercial… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/Nemotron-Instruction-Following-Chat-v1.text100K<n<1M0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.