datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NSFW_Chat_Dataset
💕 Spicy AI GF Chat Dataset 🔥
🚨 18+ Only! NSFW & Spicy Content Ahead 🚨
Hey there, AI enthusiasts and romance lovers! 😏 Welcome to the Spicy AI GF Chat Dataset, the ultimate dataset designed to bring your AI waifu to life! 💖 If you've ever dreamed of building an AI that responds like your virtual girlfriend, THIS is the dataset for you.
📜 What’s Inside?
This dataset features two columns:
input → Boyfriend’s dialogue (aka what YOU say 😉)
output →… See the full description on the dataset page: https://huggingface.co/datasets/utsavm/NSFW_Chat_Dataset.All-CVE-Chat-MultiTurn-1999-2025-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.Traditional_Chinese_roleplay_chat_Dataset
Traditional_Chinese_roleplay_chat_Dataset
這個資料集是以繁體中文為主,將各種由ChatGPT生成與極小部分個人撰寫的對話內容整理為alpaca dataset format的格式
以一層一層堆疊的方式,將一則對話紀錄拆成數筆資料(共約1000則對話),在幾次嘗試性的訓練中能夠讓llama2重現原本英文那種很活躍的對話風格,並且能夠維持善於扮演各種角色的能力
目前個人有以這個資料集製作一個lora
2023/09/07 更新
為資料集加入一些中英翻譯的句子,以期AI能以更好的文字去描寫他的動作,並增加了一些與食物有關的對話,希望能降低AI生出奇怪食物名的機率
chat_dataset_bilibilidata borrows from https://github.com/linyiLYi/bilibot
fitness-chat-prompt-completion-datasetai-girlfriend-chat-datasetmedical_dataset_chatAll-CVE-Chat-MultiTurn-1999-2025-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ukcli/All-CVE-Chat-MultiTurn-1999-2025-Dataset.legal-chat-sft-dataset
Thai Legal Chat SFT Dataset (CoT & Hybrid RAG)
ชุดข้อมูลสำหรับการทำ Instruction Fine-Tuning (SFT) เพื่อสร้าง AI ผู้ช่วยนักกฎหมายไทยที่มีความสามารถในการคิดวิเคราะห์แบบเป็นขั้นตอน (Chain-of-Thought) และมีความรู้กฎหมายที่ทันสมัยจากการใช้ Hybrid RAG (Retrieval-Augmented Generation)
Dataset Summary
ชุดข้อมูลนี้ถูกสร้างขึ้นแบบสังเคราะห์ (Synthetic Data) โดยใช้โมเดลภาษาขนาดใหญ่ (LLM) ตระกูล Qwen (27B+) บนขุมพลัง AMD MI300X ผ่านระบบ vLLM Monster Engine… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/legal-chat-sft-dataset.pashto-reasoning-chat-dataset
Pashto Reasoning Chat Dataset
A specialized chain-of-thought and multi-turn instruction dataset designed for training culturally grounded, sociologically aware, and reasoning-capable conversational AI agents in Pashto.
📊 Dataset Structure
Each sample in the dataset follows a structured conversational and reasoning format to support advanced alignment and chain-of-thought capabilities:
system: Fixed persona instructions (e.g., sociological context, cultural… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-chat-dataset.Alpha_Chat_Style_Dataset
🦾 Alpha Chat Style Dataset | darkknight25
Inject dominance, charm, and precision into your LLMs.
Crafted by Sunny Thakur, this dataset is designed to train conversational agents that speak like a leader, think like a tactician, and respond like a professional.
“Control the tone. Command the room. Every word should land like a calculated move.” – Alpha Protocol
🎯 Purpose
This dataset enables large language models—like Mixtral 8x7B Instruct—to adopt a bold… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Alpha_Chat_Style_Dataset.Vietnamese-Legal-Chat-Dataset
VLSP 2025 Vietnamese Legal Dataset
This dataset is part of the VLSP 2025 Legal SLM Challenge, designed to evaluate and train large language models on Vietnamese legal reasoning tasks.
It follows the ShareGPT conversation format, enabling supervised fine-tuning (SFT) of chat-based LLMs such as Qwen3-4B-Vietnamese-Legal-Chat.
📘 Dataset Description
Tasks Included: Multiple Choice, Natural Language Inference (NLI), and Syllogistic Legal Reasoning.
Format: Follows the… See the full description on the dataset page: https://huggingface.co/datasets/luanngo/Vietnamese-Legal-Chat-Dataset.basic-chat-model-datasetUniversal-Chat-SFT-Dataset
Universal-Chat-SFT-Dataset
A large-scale multi-turn conversational dataset designed for supervised fine-tuning (SFT) of modern Large Language Models (LLMs).
The dataset follows the OpenAI-style chat format with structured messages, making it directly compatible with most modern LLM training frameworks including Hugging Face Transformers, TRL, Axolotl, Unsloth, LlamaFactory, and custom fine-tuning pipelines.
Features
✅ Multi-turn conversations
✅ OpenAI-compatible… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Chakraborty/Universal-Chat-SFT-Dataset.Liquid-Urdu-Reasoning-Chat-Dataset
Liquid Urdu Reasoning Chat Dataset
This dataset contains high-quality, synthetically generated Urdu conversational and reasoning data designed for Supervised Fine-Tuning (SFT) of Small Language Models (SLMs), specifically optimized for reasoning models like those run via llama.cpp using DeepSeek-style reasoning formats.
Dataset Structure
The dataset is formatted in JSONL where each line contains chat history including system prompts, user queries, model… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Liquid-Urdu-Reasoning-Chat-Dataset.Sindhi-Reasoning-Chat-Dataset
# 📘 **Sindhi‑Reasoning‑Chat‑Dataset**
A high‑quality, reasoning‑focused, instruction‑tuned conversational dataset for Sindhi (سنڌي), designed to support modern NLP research and LLM training for one of South Asia’s most underrepresented languages.
---
## 🧠 **Dataset Summary**
Sindhi is a **low‑resource South Asian language** spoken by **over 30 million people**, primarily in Sindh, Pakistan. Despite its rich cultural and literary heritage, Sindhi has **very limited publicly available… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Sindhi-Reasoning-Chat-Dataset.dataset-viber-chat-generation-preference-inference-endpoints-battle
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-chat-generation-preference-inference-endpoints-battle.legal-chat-sft-dataset-thaiDevanagari-Ecommerce-fomatted-for-llama2-chat-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.ZERO-lm-dataset-chat-smallA combination of multiple open datasets combined into one for training simple lms.
Traditional_Chinese_roleplay_chat_Dataset
Traditional_Chinese_roleplay_chat_Dataset
這個資料集是以繁體中文為主,將各種由ChatGPT生成與極小部分個人撰寫的對話內容整理為alpaca dataset format的格式
以一層一層堆疊的方式,將一則對話紀錄拆成數筆資料(共約1000則對話),在幾次嘗試性的訓練中能夠讓llama2重現原本英文那種很活躍的對話風格,並且能夠維持善於扮演各種角色的能力
目前個人有以這個資料集製作一個lora
2023/09/07 更新
為資料集加入一些中英翻譯的句子,以期AI能以更好的文字去描寫他的動作,並增加了一些與食物有關的對話,希望能降低AI生出奇怪食物名的機率
query-chat-datasetchat-dataset-baselineA mirror of hikariming/chat-dataset-baseline
@misc{chat-dataset-baseline,
author = {Liu, Beiming and Huang, Kunhao and Jiao, Lihua and He, Yuchen and Zhang, Ruiqin and Liang, Yuan and Wang, Yingshan},
title = {chat-dataset-baseline},
year = {2023},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/hikariming/alpaca_chinese_dataset}},
}
legal-chat-sft-dataset2haven-chat-v3-dpo-datasetasena_Chat_Dataset_tr
Turkish Chat Dataset 🇹🇷
Dataset Özeti
Turkish Chat Dataset, Google Gemini 2.5 Flash kullanılarak özel olarak üretilmiş ve çok katmanlı kalite filtreleme süreçlerinden geçirilmiş, Türkçe için kapsamlı çok-turlu konuşma veri setlerinden biridir. 150,000 premium kalite diyalog örneği içeren bu dataset, doğal ve akıcı Türkçe konuşma AI'ları geliştirmek için optimize edilmiştir.
🎯 Ne Farklı Kılıyor?
Premium AI Üretimi: Google'ın en gelişmiş Gemini 2.5 Flash… See the full description on the dataset page: https://huggingface.co/datasets/limeXx/asena_Chat_Dataset_tr.albert-camus-chat-style-chat-dataset
Albert Camus Chat + Style Dataset
This dataset contains the training data used for the Ministral Camus project.
Structure
phase1/style-train.jsonl
Phase 1 style pretraining dataset.
Format: {"text": "..."}
phase2/chat-pairs-corpus-final-clean.jsonl
Phase 2 chat dataset from corpus-derived pairs.
Format: {"messages": [{"role": "system"|"user"|"assistant", "content": "..."}, ...]}
phase2/chat-pairs-light-boost-clean.jsonl
Additional Phase 2 chat pairs for… See the full description on the dataset page: https://huggingface.co/datasets/freddm/albert-camus-chat-style-chat-dataset.fitness-chat-prompt-completion-datasetabuse_research_chat_dataset_v2llama2_chat_datasetformat
