datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JobCCC-Conversational-Job-Recommendation-Bangladesh
JobCCC: A Conversational Code-Mixed Corpus for Job Recommendation in Bangladesh
Dataset Creators
Authors: Md. Arman Hossain, Mubashir Jawad, Fariha Khandakar Moon, and Sonia Binte Siraj
Supervisor: Dr. Nafis Sadeq
Institution: Department of Computer Science & Engineering, East West University
Dataset Summary
JobCCC (Conversational Code-Mixed Corpus) is a multi-turn conversational benchmark and job recommendation dataset tailored for the… See the full description on the dataset page: https://huggingface.co/datasets/Armans33115/JobCCC-Conversational-Job-Recommendation-Bangladesh.authentic-pre1930-sft-conversational
Pre-1930 Public Domain SFT Dataset
A supervised fine-tuning (SFT) dataset derived from 27 public-domain educational texts published before 1930, sourced from the Internet Archive. The texts span a wide range of 19th and early 20th century disciplines — natural science, history, law, philosophy, grammar, and more — and were written in a question-and-answer catechism format, making them naturally suited for instruction tuning.
Dataset Summary
Metric
Count… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/authentic-pre1930-sft-conversational.Saudi-Arabic-Alzheimers-Conversational-Dataset-Parameterized
Saudi Arabic Alzheimer's Patient QA Dataset (Conversational)
Overview
This dataset contains parameterized question-answer pairs designed for conversational AI assistants supporting Alzheimer's patients. The questions are written in the Saudi Arabic dialect and cover common memory-related interactions.
Features
Saudi Arabic dialect
Parameterized answers
Alzheimer's memory support
Conversational QA
RAG-ready
Language
Arabic (Saudi… See the full description on the dataset page: https://huggingface.co/datasets/ShahadAljohani/Saudi-Arabic-Alzheimers-Conversational-Dataset-Parameterized.harry_potter_conversational
Dataset: Harry potter conversational text corpus
Dataset Details
This corpus contains conversational data in text format
Usage
text classification
token classification
question answering
Language
en
License
apache 2.0
Conversational-Cancer-Lung-Detection
Conversational Cancer Lung Detection Dataset
This dataset, Conversational Cancer Lung Detection, is a conversationally structured dataset derived from the original Lung Cancer Detection dataset by Jillani Soft Tech on Kaggle. It has been transformed to simulate medical records in a conversational format, enabling AI applications to interact in a question-answer style format about lung cancer detection.
Dataset Overview
The Conversational Cancer Lung Detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/BrokenSoul/Conversational-Cancer-Lung-Detection.cole-conversational-corpus
Cole Multimodal Conversational Corpus
Real human-AI dialogue collected from a production AI assistant deployed across 7 messaging platforms. This is not synthetic data - these are real conversations where users send text, photos, voice messages, video, and documents alongside natural conversation.
What makes this dataset unique
Multimodal: Text, images, audio, video, and documents in conversation context - not isolated media files
Cross-channel: Same users move… See the full description on the dataset page: https://huggingface.co/datasets/huggingzetro/cole-conversational-corpus.Mental_health_data_conversationalDataset Summary
This dataset comprises a collection of questions and answers derived from two online platforms specializing in counseling and therapy. The questions address a variety of mental health topics, while the answers are written by qualified psychologists. It is designed to facilitate the fine-tuning of language models for enhanced capability in offering mental health advice.
Supported Tasks and Leaderboards
The primary task supported by this dataset is text generation, specifically… See the full description on the dataset page: https://huggingface.co/datasets/Ayansk11/Mental_health_data_conversational.prop-trading-qa-conversational-ai
Prop Trading Q&A Dataset for Conversational AI
Description
This dataset contains 200+ curated question-answer pairs covering the domain of proprietary (prop) trading firms. It is designed to serve as training and retrieval data for building AI assistants, chatbots, and educational tools focused on prop trading knowledge.
Each entry consists of a natural-language question paired with a detailed, factual answer. The data spans ten thematic categories ranging from… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/prop-trading-qa-conversational-ai.nietzsche_conversational_pashto
Nietzsche Conversational Dataset (Pashto)
This dataset provides conversational instruction-tuning data based on Friedrich Nietzsche's philosophical works, specifically tailored for conversational AI and large language models. It includes detailed internal reasoning steps (<thought> tags) in Pashto, analyzing the philosophical questions, grounding them in specific textual sections, and outlining the reasoning approach before delivering the response.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/nietzsche_conversational_pashto.sft-conversational_datasetQuestion – Answer DatasetThe dataset contains 400 queries from two domains: Current Affairs and Creative Writing. It serves as a versatile resource for Natural Language Processing (NLP) tasks, including text classification, information retrieval, and model training.
Data attributes:
Query: The user-generated question. Data type: string.
Answer: The response provided by a team of writers and editors in markdown format, containing information related to the query.
Citations: Up to 4 credible… See the full description on the dataset page: https://huggingface.co/datasets/SoftAge-AI/sft-conversational_dataset.
