CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenGVLab /InternVL-Chat-V1-2-SFT-Data Data Card for InternVL-Chat-V1-2-SFT-Data Overview Inspired by LLaVA-NeXT, we adopted a data-efficient SFT strategy to train InternVL-Chat-V1-2, utilizing approximately 1.2M of visual instruction tuning samples in total, all of which are fully open-source. In a macro sense, we build upon ShareGPT-4V and additionally integrate LLaVA-ZH, DVQA, ChartQA, AI2D, DocVQA, GeoQA+, and SynthDoG-EN. Most of the data remains consistent with LLaVA-NeXT. Citation If you use… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-Chat-V1-2-SFT-Data.imagevisual-question-answering100K<n<1M29 likes761 downloads2y agoHugging Face02ChatTSRepo /ChatTS-Training-Dataset ChatTS-Training Data This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model. Datasets align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256. align_random: Alignment training dataset with random sequence lengths between 64 and 1024. sft: SFT dataset generated with Time Series Evol-Instruct. ift: Instruction following dataset. dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.textquestion-answering100K<n<1M14 likes652 downloads1y agoHugging Face03utsavm /NSFW_Chat_Dataset 💕 Spicy AI GF Chat Dataset 🔥 🚨 18+ Only! NSFW & Spicy Content Ahead 🚨 Hey there, AI enthusiasts and romance lovers! 😏 Welcome to the Spicy AI GF Chat Dataset, the ultimate dataset designed to bring your AI waifu to life! 💖 If you've ever dreamed of building an AI that responds like your virtual girlfriend, THIS is the dataset for you. 📜 What’s Inside? This dataset features two columns: input → Boyfriend’s dialogue (aka what YOU say 😉) output →… See the full description on the dataset page: https://huggingface.co/datasets/utsavm/NSFW_Chat_Dataset.textquestion-answering1K<n<10K13 likes538 downloads2y agoHugging Face04avaliev /chat_doctorThis dataset was formed from the three data sources from the ChatDoctor work. 100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED 10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED 5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually) data sample: {'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/chat_doctor.textquestion-answering100K<n<1M15 likes228 downloads3y agoHugging Face055CD-AI /Vietnamese-Multi-turn-Chat-Alpacatextquestion-answering10K<n<100K29 likes218 downloads2y agoHugging Face06sunzeyeah /chinese_chatgpt_corpus Dataset Card for chinese_chatgpt_corpus Dataset Summary This repo collects chinese corpus for Supervised Finetuning (SFT) and Reinforcement Learning From Human Feedback (RLHF). Supported Tasks and Leaderboards More Information Needed Languages Chinese Dataset Structure Data Instances train_data_external_v1.jsonl Size of downloaded dataset files: 5.04 GB Size of the generated dataset: 0 GB… See the full description on the dataset page: https://huggingface.co/datasets/sunzeyeah/chinese_chatgpt_corpus.texttext-generation1M<n<10M88 likes182 downloads4y agoHugging Face07Naholav /cukurova_university_chatbot Çukurova University Computer Engineering Chatbot Dataset 📊 Dataset Overview This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information. 🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.textquestion-answering10K<n<100K0 likes138 downloads1y agoHugging Face08random-long-int /Java_method2test_chatml Java Method to Test ChatML This dataset is based on the methods2test dataset from Microsoft. It follows the ChatML template format: [{'role': '', 'content': ''}, {...}]. Originally, methods2test contains only Java methods at different levels of granularity along with their corresponding test cases. The different focal method segmentations are illustrated here: To simulate a conversation between a Java developer and an AI assistant, I introduce two key parameters: The prompt… See the full description on the dataset page: https://huggingface.co/datasets/random-long-int/Java_method2test_chatml.textquestion-answering100K<n<1M2 likes126 downloads2y agoHugging Face09berwart /SE-Chatting.en SE.02 Dataset Hello, welcome to the official main dataset of SE.02 that's always getting updated, make sure to like to help us a lot. this dataset contains pretty much everything from math to isk what to put but pretty much anything you think an ai can say. anyways this is our biggest dataset yet, and out first one that I don't even know how I managed to make this. you can use it to train your own ai if you want. textquestion-answering10M<n<100M5 likes120 downloads2y agoHugging Face10Chinese-Vicuna /instruct_chat_50k.jsonlinstruct_chat_50k.jsonl which is composed of 30k Chinese sharegpt dataset and 20k alpaca-instruction-Chinese-dataset textquestion-answering10K<n<100K44 likes102 downloads3y agoHugging Face11SustcZhangYX /ChatEnv ChatEnv: A Domain-Specific Instruction Dataset for Environmental Science ChatEnv is a large-scale, domain-specific instruction dataset designed to enhance large language models (LLMs) for environmental science tasks. This dataset is an integral part of the EnvGPT framework, supporting fine-tuning and evaluation processes by providing a diverse and high-quality set of instructions tailored to the unique demands of environmental science research and applications. 📃 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SustcZhangYX/ChatEnv.documentquestion-answering100K<n<1M2 likes86 downloads1y agoHugging Face12Ayansk11 /Chatkanoontabularquestion-answering100K<n<1M0 likes84 downloads2y agoHugging Face13Renicames /turkish-law-chatbot MindLaw için Hukuk Veri Seti Bu veri seti, MindLaw modelinin eğitimi için oluşturulmuş olup, Türkçe hukuk alanına özgü metinlerden derlenmiştir. Veri seti, anayasanın sunduğu içeriklerden ve anayasayı açıklayan hukuki metinlerden oluşmaktadır. Ayrıca, bireylerin avukatlara sıkça yönlendirebilecekleri sorular formatında düzenlenmiş hukuki sorular ve cevapları da içermektedir. Veri Seti İçeriği Anayasa Metinleri: Türkiye Cumhuriyeti Anayasası'nın çeşitli maddeleri ve… See the full description on the dataset page: https://huggingface.co/datasets/Renicames/turkish-law-chatbot.textquestion-answering10K<n<100K31 likes80 downloads2y agoHugging Face14Jackrong /LogicMind-Chat-Reasoning-SFT-300K Nemotron-Post-Training-Dataset-v2-chat Dataset Card Overview 📌 This dataset contains 296,168 chat-style instruction/response samples generated by qwen-3-32b. Each record provides a user prompt, an explicit reasoning trace, and a final answer, plus precomputed length fields. The data is packaged as JSONL (one JSON object per line). Highlights Scale: 296,168 samples Category: chat (100%) Generator: qwen-3-32b (100%) Structure: problem → qwen3-reasoning → qwen3-solution… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/LogicMind-Chat-Reasoning-SFT-300K.tabularquestion-answering100K<n<1M10 likes80 downloads8mo agoHugging Face15oddadmix /arabic-rag-chat-8k-eval arabic-rag-chat-8k-eval Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models: the test split, every model's raw replies, every judge verdict, and the rendered report for each. Thirteen judged models, all scored on the same 1,651 prompts by the same judge at temperature 0.0, so the comparison below is like-for-like and can be recomputed offline without a GPU or a judge server. This is the measurement half of oddadmix/100M-8192-Nawah-dsv4; the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.tabularquestion-answeringn<1K0 likes72 downloads1mo agoHugging Face16mesolitica /chatgpt4-commonsense-qa Synthetic CommonSense Generated using ChatGPT4, originally from https://huggingface.co/datasets/commonsense_qa Notebook at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/chatgpt4-commonsense synthetic-commonsense.jsonl, 36332 rows, 7.34 MB. Example data {'question': '1. Seseorang yang bersara mungkin perlu kembali bekerja jika mereka apa?\n A. mempunyai hutang\n B. mencari pendapatan\n C. meninggalkan pekerjaan\n D. memerlukan… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt4-commonsense-qa.textquestion-answering10K<n<100K1 likes71 downloads3y agoHugging Face17nyuuzyou /chatgpt-in-russia-qa Dataset Card for чатгпт-в-россии.рф Dataset Summary This dataset contains question-answer pairs collected from чатгпт-в-россии.рф (meaning in English would be something like chatgpt-in-russia[.]rf), a Russian question-answering website. Each entry in the dataset represents a question asked by a user and the corresponding answer generated by an unspecified language model. The dataset contains 704,208 unique question-answer pairs covering various topics.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/chatgpt-in-russia-qa.textquestion-answering100K<n<1M15 likes65 downloads1y agoHugging Face18haiderkamal23 /allaM-offsec-arabic-chat-v2 Arabic Offensive Security Chat Dataset v2 High-quality category-aware bilingual Arabic/English dataset for offensive security assistants. What's New in v2 ✅ Category-aware responses: Different response structures for web vulns, DeFi, reconnaissance tools, social engineering, etc. ✅ No generic templates: Each category has specialized analysis framework ✅ No verbatim copying: Responses analyze and transform the input, not repeat it ✅ Semantic accuracy: Tools (nmap… See the full description on the dataset page: https://huggingface.co/datasets/haiderkamal23/allaM-offsec-arabic-chat-v2.textquestion-answering10K<n<100K0 likes56 downloads10mo agoHugging Face19Jackrong /Qwen3-235B-A22B-Instruct-2507-Distilled-chat Qwen3-235B-A22B-Instruct-2507-Distilled-chat📚 Curated/Funded/Shared by: [Jack Rong] Language(s): English (major), Chinese, Русский, 한국어, 日本語, others License: [apache-2.0] Distilled Model: 🏆Qwen/Qwen3-235B-A22B-Instruct-2507 Qwen3-235B-A22B-Instruct-2507 Benchmarks📊 Introduction: The objectives of this project are: Focus on chat capabilities (excluding CoT), covering cross-lingual real-world Q&A/explanation/generation; Utilize… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Qwen3-235B-A22B-Instruct-2507-Distilled-chat.texttable-question-answering1K<n<10K6 likes54 downloads1y agoHugging Face20oddadmix /arabic-rag-chat-30K Arabic multi-turn RAG customer-support conversations (31,294 conversations) Synthetic Modern Standard Arabic customer-support conversations for training small Arabic RAG assistants. The bulk was distilled from gemini-3.1-flash-lite via the Batch API; a first 2.4% came from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM server before the run was moved off-GPU. Both teachers were given the same prompts and the same validator. Each row is one conversation of 1-5 rounds over one… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-30K.tabularquestion-answering10K<n<100K0 likes51 downloads1mo agoHugging Face21lianghsun /tw-bar-examination-2020-chat Dataset Card for tw-bar-examination-2020-chat tw-bar-examination-2020-chat 是一個中華民國 2020 年律師考試選擇題之 Alpaca 格式微調資料集,合計 299 題(train 269、test 30)。每題包含統一提示語、題目與四個選項,以及正確答案字母,適用於微調繁體中文語言模型於台灣法律選擇題作答任務。 Dataset Details Dataset Description 本資料集源自 Jamie0510/taiwan-law-exam 中之 2020 年律師考試題目,整合其四大類科後進行後處理:去除欄位缺失之題目,並統一轉為 Alpaca 三欄格式(instruction / input / output)。每題之 instruction 欄為固定提示語「請在下列的單一選擇題中,選出正確的答案,並且只回答 A, B, C, D 其中一個字代表正確答案」。 本資料集作為 SFT 訓練素材設計,建議與… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-bar-examination-2020-chat.textquestion-answeringn<1K3 likes50 downloads5mo agoHugging Face22Jackrong /Chinese-DeepSeek-V3.2-Exp-chat-example deepseek/deepseek-v3.2-exp (6.6K) 中文数据集样本 一、前言 本报告基于 deepseek/deepseek-v3.2-exp 模型(官方 API,8K 上下文窗口)进行数据集评测与可视化展示。测试数据集共包含 6,655 轮对话,语言覆盖以中文为主,辅以部分混合语种及非中文输入。本次报告旨在总结模型的对话特征、输入输出长度分布及上下文预算消耗情况,并为后续应用和优化提供参考。 二、数据与方法 数据来源:用户构建的 6,655 轮真实中文对话样本。 估算方法: 中文字符近似为 1 Token; 英文 4 字符 ≈ 1 Token; 用于规模与上下文预算对比,而非精确 Token 计数。 统计维度: 平均 Prompt/Output 长度(字符与估算 Token); 总 Token 占上下文窗口比例; 语言分布(Prompt 语言类型); 对话长度分布(用户提问、助手回答、总对话长度)。 三、总体结果 1. 样本概况… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-DeepSeek-V3.2-Exp-chat-example.tabularquestion-answering1K<n<10K5 likes47 downloads1y agoHugging Face23monsoon-nlp /asknyc-chatassistant-formatQuestions from Reddit.com/r/AskNYC, downloaded from PushShift, filtered to direct responses from humans, where the post net score is >= 3. Collected one month of posts from each year 2015-2019 (i.e. no content from July 2019 onward) Adapted from the CSV used to fine-tune https://huggingface.co/monsoon-nlp/gpt-nyc Blog about the original model: https://medium.com/geekculture/gpt-nyc-part-1-9cb698b2e3d textquestion-answering10K<n<100K0 likes46 downloads2y agoHugging Face24Sr523 /big-red-bark-chat-evaluation Big Red Bark Chat Q&A Dataset Dataset Description This dataset contains 12,385 question-and-answer pairs collected from Big Red Bark Chat, an innovative AI assistant developed at Cornell University that answers questions about dog health (as well as other animal species). While it does not replace professional veterinary advice, it serves as a valuable starting point by searching trusted sources. Big Red Bark Chat is designed to provide quick and reliable answers… See the full description on the dataset page: https://huggingface.co/datasets/Sr523/big-red-bark-chat-evaluation.textquestion-answering10K<n<100K0 likes46 downloads3mo agoHugging Face25Fredithefish /Instruction-Tuning-with-GPT-4-RedPajama-Chat Instruction Tuning with GPT 4 RedPajama-Chat This dataset has been converted from the Instruction-Tuning-with-GPT-4 dataset for the purpose of fine-tuning the RedPajama-INCITE-Chat-3B-v1 model. About Instruction-Tuning-with-GPT-4 English Instruction-Following Data generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. Usage and License Notices The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Instruction-Tuning-with-GPT-4-RedPajama-Chat.textquestion-answering10K<n<100K6 likes45 downloads3y agoHugging Face26Phonsiri /legal-chat-sft-dataset Thai Legal Chat SFT Dataset (CoT & Hybrid RAG) ชุดข้อมูลสำหรับการทำ Instruction Fine-Tuning (SFT) เพื่อสร้าง AI ผู้ช่วยนักกฎหมายไทยที่มีความสามารถในการคิดวิเคราะห์แบบเป็นขั้นตอน (Chain-of-Thought) และมีความรู้กฎหมายที่ทันสมัยจากการใช้ Hybrid RAG (Retrieval-Augmented Generation) Dataset Summary ชุดข้อมูลนี้ถูกสร้างขึ้นแบบสังเคราะห์ (Synthetic Data) โดยใช้โมเดลภาษาขนาดใหญ่ (LLM) ตระกูล Qwen (27B+) บนขุมพลัง AMD MI300X ผ่านระบบ vLLM Monster Engine… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/legal-chat-sft-dataset.texttext-generation10K<n<100K1 likes45 downloads5mo agoHugging Face27ai-forever /paper_persi_chat PaperPersiChat Dataset Dataset for paper PaperPersiChat: Scientific Paper Discussion Chatbot using Transformers and Discourse Flow Management Dataset creation To construct the dataset, we used the part of Semantic Scholar Open Research Corpus [https://github.com/allenai/s2orc] as the main source of scientific publications, namely the Computer Science section. We constructed dialogues over the segments of the papers where each segment consists of a combination of several… See the full description on the dataset page: https://huggingface.co/datasets/ai-forever/paper_persi_chat.texttext-generation10K<n<100K10 likes44 downloads3y agoHugging Face28darkknight25 /Alpha_Chat_Style_Dataset 🦾 Alpha Chat Style Dataset | darkknight25 Inject dominance, charm, and precision into your LLMs. Crafted by Sunny Thakur, this dataset is designed to train conversational agents that speak like a leader, think like a tactician, and respond like a professional. “Control the tone. Command the room. Every word should land like a calculated move.” – Alpha Protocol 🎯 Purpose This dataset enables large language models—like Mixtral 8x7B Instruct—to adopt a bold… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Alpha_Chat_Style_Dataset.textquestion-answering1K<n<10K0 likes40 downloads1y agoHugging Face29MichaelAnthony /lemonseed-rl-chat-tasks lemonseed-rl-chat-tasks LemonSeed — chat-alignment RL tasks (prompt/gold single-turn). Contents rl_chat_tasks.jsonl (9852 rows) Format JSON Lines (.jsonl), one example per line. Provenance Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella. textquestion-answering1K<n<10K0 likes37 downloads1mo agoHugging Face30oddadmix /arabic-rag-chat-grpo-5K Arabic multi-turn RAG conversations — GRPO pool (5,259 conversations) The reinforcement-learning half of oddadmix/arabic-rag-chat-30K: same generator, same validator, same schema, disjoint companies. It exists so GRPO explores fresh knowledge bases instead of taking a second pass over material the SFT already memorised. conversations turns companies this pool 5,259 14,018 309 Company-disjointness is exact and verified: this pool shares zero company_id values… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-grpo-5K.tabularquestion-answering1K<n<10K1 likes36 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.