datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dailyconversationsThis dataset is synthetically generated using ChatGPT 3.5 to contain two-person multi-turn daily conversations with a various of topics (e.g.
travel, food, music, movie/TV, education, hobbies, family, sports, technology, books, etc.) Originally, this dataset is used to train
QuicktypeGPT, which is a GPT model to assist auto complete conversations.
Here is the full list of topics the conversation may cover.
simple-daily-conversations-cleaned
Dataset Card
This dataset contains a cleaned version of simple daily conversations. It comprises nearly 98K text snippets representing informal, everyday dialogue, curated and processed for various Natural Language Processing tasks.
Uses
Direct Use
This dataset is ideal for:
Training language models on informal, everyday conversational data.
Research exploring linguistic patterns in casual conversation.
Out-of-Scope Use
The dataset may not… See the full description on the dataset page: https://huggingface.co/datasets/aarohanverma/simple-daily-conversations-cleaned.daily-conversation-batch-01-v5-clean
daily-conversation-batch-01-v5-clean
Strictly filtered Arabic multi-turn conversation subset.
Summary
Source set: previously cleaned keep pool (A set) from arabic-daily-batch01-v5-5k-data.jsonl
Second-pass reviewed records so far: 1311
Kept after strict review: 311
Dropped after strict review: 1000
Keep rate over reviewed subset: 23.72%
Strict review bar
Each kept sample was screened on all of:
assistant hidden thinking quality
visible assistant reply… See the full description on the dataset page: https://huggingface.co/datasets/Jianshu001/daily-conversation-batch-01-v5-clean.Translate-Khmer-Daily-Casual-Conversation
Translate Khmer Daily Casual Conversation
Welcome to the SeyhaLite collection. This dataset has been meticulously curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) and Translation systems.
Project Vision
I hope this dataset helps your project succeed! Whether you are building a translation tool, an AI assistant, or conducting research, this data is designed to provide clear and accurate information about everyday interactions… See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Translate-Khmer-Daily-Casual-Conversation.daily_conversationsconversation_daily_multiturnDaily_Conversation_Hinglisht1_daily_conversations_v3english-daily-conversationt1_daily_conversations_v2t1_daily_conversationssimple-daily-conversations-cleaned-ckb
