datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NSFW_Chat_Dataset
💕 Spicy AI GF Chat Dataset 🔥
🚨 18+ Only! NSFW & Spicy Content Ahead 🚨
Hey there, AI enthusiasts and romance lovers! 😏 Welcome to the Spicy AI GF Chat Dataset, the ultimate dataset designed to bring your AI waifu to life! 💖 If you've ever dreamed of building an AI that responds like your virtual girlfriend, THIS is the dataset for you.
📜 What’s Inside?
This dataset features two columns:
input → Boyfriend’s dialogue (aka what YOU say 😉)
output →… See the full description on the dataset page: https://huggingface.co/datasets/utsavm/NSFW_Chat_Dataset.dutch_chat_datasets
Dataset Card for "dutch_chat_datasets"
This dataset is a merge of the following datasets. See their pages for licensing, usage, creation, and citation information.
https://huggingface.co/datasets/BramVanroy/dolly-15k-dutch
https://huggingface.co/datasets/BramVanroy/alpaca-cleaned-dutch-baize
https://huggingface.co/datasets/BramVanroy/stackoverflow-chat-dutch
https://huggingface.co/datasets/BramVanroy/quora-chat-dutch
They are reformatted for easier, consistent processing in… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/dutch_chat_datasets.healthcare-chat-dataset-jsonl
Healthcare Chat Dataset (JSONL Format)
This dataset contains 41 healthcare-related conversational exchanges in ChatML format, designed for training conversational AI models for medical assistance and healthcare guidance.
Dataset Structure
The dataset is provided as a JSONL file where each line contains a JSON object with:
text: A complete conversation in ChatML format with system, user, and assistant messages
ChatML Format Structure
Each conversation follows… See the full description on the dataset page: https://huggingface.co/datasets/adrianf12/healthcare-chat-dataset-jsonl.legal-chat-sft-dataset
Thai Legal Chat SFT Dataset (CoT & Hybrid RAG)
ชุดข้อมูลสำหรับการทำ Instruction Fine-Tuning (SFT) เพื่อสร้าง AI ผู้ช่วยนักกฎหมายไทยที่มีความสามารถในการคิดวิเคราะห์แบบเป็นขั้นตอน (Chain-of-Thought) และมีความรู้กฎหมายที่ทันสมัยจากการใช้ Hybrid RAG (Retrieval-Augmented Generation)
Dataset Summary
ชุดข้อมูลนี้ถูกสร้างขึ้นแบบสังเคราะห์ (Synthetic Data) โดยใช้โมเดลภาษาขนาดใหญ่ (LLM) ตระกูล Qwen (27B+) บนขุมพลัง AMD MI300X ผ่านระบบ vLLM Monster Engine… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/legal-chat-sft-dataset.Alpha_Chat_Style_Dataset
🦾 Alpha Chat Style Dataset | darkknight25
Inject dominance, charm, and precision into your LLMs.
Crafted by Sunny Thakur, this dataset is designed to train conversational agents that speak like a leader, think like a tactician, and respond like a professional.
“Control the tone. Command the room. Every word should land like a calculated move.” – Alpha Protocol
🎯 Purpose
This dataset enables large language models—like Mixtral 8x7B Instruct—to adopt a bold… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Alpha_Chat_Style_Dataset.italian-open-sft-chat-dataset
Italian Open SFT Chat Dataset
An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records.
This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.basic-chat-model-datasetchat_dataset
Dataset Card for Llama Mode Dataset
Dataset Details
Dataset Description
The Llama Mode dataset is a specialized educational dataset designed to facilitate the development of AI models for interactive learning with a focus on engaging students with diverse needs, including those in Special Educational Needs (SEN) education. This dataset contains prompts and responses paired with educational strategies and their applications, particularly structured to assist and… See the full description on the dataset page: https://huggingface.co/datasets/AllyArc/chat_dataset.Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.chat_dataset-2
