CoolFace
Datasetpublic

AIxBlock/Human-to-machine-Japanese-audio-call-center-conversations

Dataset Card for Japanese audio call center human to machine conversations This dataset contains synthetic audio conversations in Japanese between human customers and machine agents, simulating real-world call center scenarios Dataset Details Dataset Description Curated by: AIxBlock (aixblock.io) Funded by [optional]: AIxBlock (aixblock.io) Shared by [optional]: AIxBlock (aixblock.io) Language(s) (NLP): Japanese License: Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Human-to-machine-Japanese-audio-call-center-conversations.

sourceHugging Facecc-by-nc-4.0updated 1y agoView on Hugging Face
2likes61downloads
Dataset Card

Dataset Card for Japanese audio call center human to machine conversations

<!-- Provide a quick summary of the dataset. -->

This dataset contains synthetic audio conversations in Japanese between human customers and machine agents, simulating real-world call center scenarios

Dataset Details

Dataset Description

<!-- Provide a longer summary of what this dataset is. -->

  • —Curated by: AIxBlock (aixblock.io)
  • —Funded by [optional]: AIxBlock (aixblock.io)
  • —Shared by [optional]: AIxBlock (aixblock.io)
  • —Language(s) (NLP): Japanese
  • —License: Creative Commons Attribution Non Commercial 4.0
  • —This dataset contains synthetic audio conversations between human customers and machine agents, simulating real-world call center scenarios. Each conversation is designed to reflect natural interaction patterns across a variety of customer service topics.

🗣️ Participants: Human speakers are native Japanese users of diverse ages and genders, ensuring a wide range of speaking styles and acoustic characteristics.

🤖 Dialogue Type: Human-to-machine interactions where the machine agent follows scripted logic to handle queries.

🎧 Audio Format: High-quality stereo audio recordings, wav, suitable for speech recognition, speaker diarization, and conversational AI research.

Dataset Sources [optional]

<!-- Provide the basic links for the dataset. -->

  • —Link to download in full (pls note that data is heavy so we cant upload all files directly to HF): https://drive.google.com/drive/folders/1NiqwmP7kT9lF7IKIlv5Mo5b7x2EKCn?usp=sharing

Uses

<!-- Address questions around how the dataset is intended to be used. -->

Direct Use

<!-- This section describes suitable use cases for the dataset. -->

You can download this dataset and use it to fine-tune your models.

Out-of-Scope Use

<!-- This section addresses misuse, malicious use, and uses that the dataset will not work well for. -->

You are not allowed to download this dataset and re-sell to someome.

Personal and Sensitive Information

<!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. -->

All PII info in conversations are made up, but following real patterns.

Citation [optional]

<!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. -->

More Information [optional]

More info needed, pls contact us via our discord channel: https://discord.gg/nePjg9g5v6

Dataset Card Authors [optional]

Brought to you by AIxBlock, a decentralized AI development and workflow automation platform

Dataset Card Contact

Discord: https://discord.gg/nePjg9g5v6 Please give us a follow on HF if you want to access our newly released dataset. Follow us on Github to access some rare dataset and our sourcecode: https://github.com/AIxBlock-2023/aixblock-ai-dev-platform-public