datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cukurova_university_chatbot
Çukurova University Computer Engineering Chatbot Dataset
📊 Dataset Overview
This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information.
🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.turkish-law-chatbot
MindLaw için Hukuk Veri Seti
Bu veri seti, MindLaw modelinin eğitimi için oluşturulmuş olup, Türkçe hukuk alanına özgü metinlerden derlenmiştir. Veri seti, anayasanın sunduğu içeriklerden ve anayasayı açıklayan hukuki metinlerden oluşmaktadır. Ayrıca, bireylerin avukatlara sıkça yönlendirebilecekleri sorular formatında düzenlenmiş hukuki sorular ve cevapları da içermektedir.
Veri Seti İçeriği
Anayasa Metinleri: Türkiye Cumhuriyeti Anayasası'nın çeşitli maddeleri ve… See the full description on the dataset page: https://huggingface.co/datasets/Renicames/turkish-law-chatbot.mental_health_Chatbot
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Iamzoo/mental_health_Chatbot.chatbotai-chatbot-scrubeedata-science-chatbot
📊 Data Science Chatbot Dataset (2000 Samples)
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts.
This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a Data Science Tutor
Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.chatbotpashto-chatbot-greetings
Pashto Chatbot Greetings Dataset
Multi-turn chatbot greeting configurations translated from Italian into Pashto (ps).
Details
Total Records: 8,846
Format: Streaming JSON Lines (.jsonl)
Maintained by: nassimjp
