CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-SFT-Instruction-Following-Chat-v3 Dataset Description: The Nemotron-Instruction-Following-Chat-v3 dataset is designed to strengthen multi-turn, interactive capabilities, including open-ended chat and precise instruction following. The chat subset uses human written prompts from sources like lmarena, lmsys, and wildchat as seed prompts. Responses are generated with GLM-5. Multiple responses are sampled from the model and the best response as judged by pairwise comparisons using Qwen3-Nemotron-235B-A22B-GenRM-2603… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.texttext-generation100K<n<1M20 likes5.1k downloads4mo agoHugging Face02nvidia /Nemotron-SFT-Instruction-Following-Chat-v2 Dataset Description: The Nemotron-Instruction-Following-Chat-v2 dataset is designed to broadly strengthen the model’s interactive capabilities, including open-ended chat and precise instruction following.The dataset is a refreshed version of Nemotron-Instruction-Following-Chat-v1 with synthetic dialogues generated from Kimi-K2-Thinking, GLM-4.6, Qwen3-235B-A22B-Thinking-2507, GPT-OSS-120b, Kimi-K2-Instruct-0905, and Qwen3-235B-A22B-Instruct-2507. This dataset is ready for commercial… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v2.text-generation31 likes4.9k downloads7mo agoHugging Face03nvidia /Nemotron-Instruction-Following-Chat-v1 Dataset Description: The Nemotron-Instruction-Following-Chat-v1 dataset is designed to broadly strengthen the model’s interactive capabilities, spanning open-ended chat, precise instruction following, and reliable structured output generation. It combines refreshed chat data from Nemotron-Post-Training-Dataset-v2 (extended to multi-turn) with synthetic dialogues produced by strong frontier models such as GPT-OSS-120B and Qwen3-235B variants. This dataset is ready for commercial… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Instruction-Following-Chat-v1.text100K<n<1M130 likes986 downloads9mo agoHugging Face04scaleinvariant /llama-3.2-1b-instruct-lmsys-chat-1m-activations Llama 3.2 1B Instruct Activations (LMSYS-Chat-1M) This dataset contains whole-model residual stream activations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Each row stores the complete residual stream across all 16 transformer layers for a single prompt — both the full-sequence activations and the final-token activations. Note: This is a subset, 8% (from 2 workers of 25) of the full dataset. The complete dataset was ~25 TB and huggingface only… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/llama-3.2-1b-instruct-lmsys-chat-1m-activations.textfeature-extraction10K<n<100K0 likes410 downloads6mo agoHugging Face05Chat-UniVi /Chat-UniVi-Instruct Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding Paper or resources for more information: [Paper] [Code] image8 likes307 downloads2y agoHugging Face06jamesdborin /Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.tabular1M<n<10M0 likes260 downloads3mo agoHugging Face07bew /lmsys-chat-1m-qwen2.5-instruct100K<n<1M0 likes237 downloads2y agoHugging Face08crystal-ai /chat-compilation-benchmark-5x-Llama-3.2-Instruct-Shuffledtext1M<n<10M0 likes207 downloads1y agoHugging Face09Hugodonotexit /chat-Instruct-en Chat-Instruct-en Aggregated English instruction-following and chat examples. Dataset Summary Hugodonotexit/chat-Instruct-en is a consolidated dataset of English-only instruction and chat samples. Each example is stored in one text field containing a <|user|> prompt and a <|assistant|> response. Upstream Sources This dataset aggregates data from: BAAI/Infinity-Instruct lmsys/lmsys-chat-1 (or the corresponding LMSYS chat release used in this collection) Before… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/chat-Instruct-en.text1M<n<10M2 likes189 downloads8mo agoHugging Face10bew /lmsys-chat-1m-qwen2-instruct-768100K<n<1M0 likes184 downloads2y agoHugging Face11bew /lmsys-chat-1m-qwen2.5-instruct-1024-context100K<n<1M0 likes141 downloads2y agoHugging Face12open-llm-leaderboard-old /details_acrastt__RedPajama-INCITE-Chat-Instruct-3B-V1 Dataset Card for Evaluation run of acrastt/RedPajama-INCITE-Chat-Instruct-3B-V1 Dataset Summary Dataset automatically created during the evaluation run of model acrastt/RedPajama-INCITE-Chat-Instruct-3B-V1 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_acrastt__RedPajama-INCITE-Chat-Instruct-3B-V1.0 likes117 downloads3y agoHugging Face13Chinese-Vicuna /instruct_chat_50k.jsonlinstruct_chat_50k.jsonl which is composed of 30k Chinese sharegpt dataset and 20k alpaca-instruction-Chinese-dataset textquestion-answering10K<n<100K44 likes111 downloads3y agoHugging Face14Chat-Error /Long-instructionstext10K<n<100K3 likes108 downloads3y agoHugging Face15Hastagaras /JF_nvidia_Nemotron-SFT-Instruction-Following-Chat-v2_reasoning_offtext100K<n<1M0 likes104 downloads6mo agoHugging Face16jamesdborin /Nemotron-Instruction-Following-Chat-and-Knowledge-prompt-only Instruction Following, Chat and Knowledge Prompt-Only This dataset combines prompt-only datasets by capability theme for distillation experiments. It contains 2,235,051 unique prompts from 3,344,905 raw rows; 1,109,854 exact canonical duplicates were removed. Rows retain the canonical prompt-extraction columns and add source_repo_id for provenance. Deduplication uses normalized system_prompt, prompt, tools, and schema_str, with the first row in manifest order retained. Original… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Instruction-Following-Chat-and-Knowledge-prompt-only.0 likes96 downloads2mo agoHugging Face17bew /lmsys-chat-1m-qwen2-instruct100K<n<1M0 likes95 downloads2y agoHugging Face18aklein4 /chat-compilation-benchmark-5x-Llama-3.2-Instruct-Shuffledtext1M<n<10M0 likes88 downloads1y agoHugging Face19amalia-llm /amalia-Nemotron-Instruction-Following-Chat-v1 AMALIA Nemotron-Instruction-Following-Chat-v1 Version of the nvidia/Nemotron-Instruction-Following-Chat-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to: Remove entries where the capability_target field was 'chat' from the chat_if split; Remove entries that reference other LLMs or research labs; Remove entries that contained the string '/imagine prompt:'; Removed the reasoning_content field; Original… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-Instruction-Following-Chat-v1.text10K<n<100K0 likes82 downloads3mo agoHugging Face20open-llm-leaderboard-old /details_Lajonbot__Llama-2-7b-chat-hf-instruct-pl-lora_unload Dataset Card for Evaluation run of Lajonbot/Llama-2-7b-chat-hf-instruct-pl-lora_unload Dataset Summary Dataset automatically created during the evaluation run of model Lajonbot/Llama-2-7b-chat-hf-instruct-pl-lora_unload on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Lajonbot__Llama-2-7b-chat-hf-instruct-pl-lora_unload.0 likes77 downloads3y agoHugging Face21open-llm-leaderboard-old /details_Fredithefish__RedPajama-INCITE-Chat-3B-Instruction-Tuning-with-GPT-4 Dataset Card for Evaluation run of Fredithefish/RedPajama-INCITE-Chat-3B-Instruction-Tuning-with-GPT-4 Dataset Summary Dataset automatically created during the evaluation run of model Fredithefish/RedPajama-INCITE-Chat-3B-Instruction-Tuning-with-GPT-4 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Fredithefish__RedPajama-INCITE-Chat-3B-Instruction-Tuning-with-GPT-4.0 likes75 downloads3y agoHugging Face22Sidsidney /Nemotron-Instruction-Following-Chat-v1 Dataset Description: The Nemotron-Instruction-Following-Chat-v1 dataset is designed to broadly strengthen the model’s interactive capabilities, spanning open-ended chat, precise instruction following, and reliable structured output generation. It combines refreshed chat data from Nemotron-Post-Training-Dataset-v2 (extended to multi-turn) with synthetic dialogues produced by strong frontier models such as GPT-OSS-120B and Qwen3-235B variants. This dataset is ready for commercial… See the full description on the dataset page: https://huggingface.co/datasets/Sidsidney/Nemotron-Instruction-Following-Chat-v1.text100K<n<1M0 likes67 downloads9mo agoHugging Face23jamesdborin /Nemotron-SFT-Instruction-Following-Chat-v3-prompt-only Nemotron-SFT-Instruction-Following-Chat-v3-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v3. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v3-prompt-only.0 likes64 downloads3mo agoHugging Face24bew /lmsys-chat-1m-qwen2.5-instruct-texttext1M<n<10M0 likes62 downloads2y agoHugging Face25Fredithefish /Instruction-Tuning-with-GPT-4-RedPajama-Chat Instruction Tuning with GPT 4 RedPajama-Chat This dataset has been converted from the Instruction-Tuning-with-GPT-4 dataset for the purpose of fine-tuning the RedPajama-INCITE-Chat-3B-v1 model. About Instruction-Tuning-with-GPT-4 English Instruction-Following Data generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. Usage and License Notices The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Instruction-Tuning-with-GPT-4-RedPajama-Chat.textquestion-answering10K<n<100K6 likes61 downloads3y agoHugging Face26aiuser3993 /Llama-Krikri-8B-Instruct-ChatCreative and vocabulary-rich (LLM-focused) Greek conversation data generated by ilsp/Llama-Krikri-8B-Instruct-GGUF for distlization purposes. textn<1K0 likes58 downloads24d agoHugging Face27Jackrong /Qwen3-235B-A22B-Instruct-2507-Distilled-chat Qwen3-235B-A22B-Instruct-2507-Distilled-chat📚 Curated/Funded/Shared by: [Jack Rong] Language(s): English (major), Chinese, Русский, 한국어, 日本語, others License: [apache-2.0] Distilled Model: 🏆Qwen/Qwen3-235B-A22B-Instruct-2507 Qwen3-235B-A22B-Instruct-2507 Benchmarks📊 Introduction: The objectives of this project are: Focus on chat capabilities (excluding CoT), covering cross-lingual real-world Q&A/explanation/generation; Utilize… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Qwen3-235B-A22B-Instruct-2507-Distilled-chat.texttable-question-answering1K<n<10K6 likes57 downloads1y agoHugging Face28AmanPriyanshu /reasoning-sft-Nemotron-Instruction-Following-Chat-v1 Nemotron Instruction Following Chat v1 (Reasoning SFT) Converted version of nvidia/Nemotron-Instruction-Following-Chat-v1, filtered to 157,595 rows where assistant responses include genuine reasoning traces (reasoning_content). Format Each row has three columns: input — list of dicts with role/content conversation turns (system, user, and prior assistant turns up to the final assistant response) response — <think> block containing the model's reasoning followed by the… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-Nemotron-Instruction-Following-Chat-v1.texttext-generation100K<n<1M0 likes53 downloads6mo agoHugging Face29O2iginal /Qwen2.5-distill-v3-instruct-dataset-base-chat-template-40B-202603220 likes52 downloads6mo agoHugging Face30OperatorSDG /Nemotron-SFT-Instruction-Following-Chat-v2 Dataset Description: The Nemotron-Instruction-Following-Chat-v2 dataset is designed to broadly strengthen the model’s interactive capabilities, including open-ended chat and precise instruction following.The dataset is a refreshed version of Nemotron-Instruction-Following-Chat-v1 with synthetic dialogues generated from Kimi-K2-Thinking, GLM-4.6, Qwen3-235B-A22B-Thinking-2507, GPT-OSS-120b, Kimi-K2-Instruct-0905, and Qwen3-235B-A22B-Instruct-2507. This dataset is ready for… See the full description on the dataset page: https://huggingface.co/datasets/OperatorSDG/Nemotron-SFT-Instruction-Following-Chat-v2.text-generation0 likes50 downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.