daily-conversation
ePark_sheng_huo_hui_hua_pian_daily_conversation
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation.dailyconversationsThis dataset is synthetically generated using ChatGPT 3.5 to contain two-person multi-turn daily conversations with a various of topics (e.g.
travel, food, music, movie/TV, education, hobbies, family, sports, technology, books, etc.) Originally, this dataset is used to train
QuicktypeGPT, which is a GPT model to assist auto complete conversations.
Here is the full list of topics the conversation may cover.
simple-daily-conversations-cleaned
Dataset Card
This dataset contains a cleaned version of simple daily conversations. It comprises nearly 98K text snippets representing informal, everyday dialogue, curated and processed for various Natural Language Processing tasks.
Uses
Direct Use
This dataset is ideal for:
Training language models on informal, everyday conversational data.
Research exploring linguistic patterns in casual conversation.
Out-of-Scope Use
The dataset may not… See the full description on the dataset page: https://huggingface.co/datasets/aarohanverma/simple-daily-conversations-cleaned.Translate-Khmer-Daily-Casual-Conversation
Translate Khmer Daily Casual Conversation
Welcome to the SeyhaLite collection. This dataset has been meticulously curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) and Translation systems.
Project Vision
I hope this dataset helps your project succeed! Whether you are building a translation tool, an AI assistant, or conducting research, this data is designed to provide clear and accurate information about everyday interactions… See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Translate-Khmer-Daily-Casual-Conversation.daily-conversation-batch-01-v5-clean
daily-conversation-batch-01-v5-clean
Strictly filtered Arabic multi-turn conversation subset.
Summary
Source set: previously cleaned keep pool (A set) from arabic-daily-batch01-v5-5k-data.jsonl
Second-pass reviewed records so far: 1311
Kept after strict review: 311
Dropped after strict review: 1000
Keep rate over reviewed subset: 23.72%
Strict review bar
Each kept sample was screened on all of:
assistant hidden thinking quality
visible assistant reply… See the full description on the dataset page: https://huggingface.co/datasets/Jianshu001/daily-conversation-batch-01-v5-clean.USA-accented-role-playing-daily-conversations-stereo
Dataset Card for Synthetic daily conversations - USA accented - stereo wav
This dataset consists of synthetic daily conversations recorded by native U.S. English speakers with authentic American accents.
Dataset Details
Dataset Description
This dataset consists of synthetic daily conversations recorded by native U.S. English speakers with authentic American accents. The dialogues are spoken spontaneously, covering a range of everyday topics such… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/USA-accented-role-playing-daily-conversations-stereo.
