CoolFace
Datasetpublic

AIxBlock/USA-accented-role-playing-daily-conversations-stereo

Dataset Card for Synthetic daily conversations - USA accented - stereo wav This dataset consists of synthetic daily conversations recorded by native U.S. English speakers with authentic American accents. Dataset Details Dataset Description This dataset consists of synthetic daily conversations recorded by native U.S. English speakers with authentic American accents. The dialogues are spoken spontaneously, covering a range of everyday topics such… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/USA-accented-role-playing-daily-conversations-stereo.

sourceHugging Facecc-by-nc-4.0updated 1y agoView on Hugging Face
5likes26downloads
Dataset Card

Dataset Card for Synthetic daily conversations - USA accented - stereo wav

<!-- Provide a quick summary of the dataset. -->

This dataset consists of synthetic daily conversations recorded by native U.S. English speakers with authentic American accents.

Dataset Details

Dataset Description

<!-- Provide a longer summary of what this dataset is. -->

This dataset consists of synthetic daily conversations recorded by native U.S. English speakers with authentic American accents. The dialogues are spoken spontaneously, covering a range of everyday topics such as scheduling, small talk, casual debates, errands, and more.

🎤 Audio Format: High-quality stereo WAV files recorded in quiet environments for clean acoustic input.

🗣️ Speakers: Diverse participants across different ages and genders, all native USA speakers.

🧾 Consent: All recordings were made with full informed consent.

💬 Speech Style: Natural, unscripted, and conversational to reflect real-life speech dynamics.

🔐 Usage Restriction: This dataset is released for the sole purpose of fine-tuning AI models. Commercial use and resale are strictly prohibited.

The dataset is ideal for training and evaluating automatic speech recognition (ASR), speaker diarization, speech synthesis, and natural language understanding (NLU) systems.

  • —Curated by:: AIxBlock
  • —Funded by [optional]: AIxBlock
  • —Shared by [optional]: AIxBlock
  • —Language(s) (NLP): English, USA accent

Dataset Sources [optional]

<!-- Provide the basic links for the dataset. -->

  • —Repository: https://github.com/AIxBlock-2023/aixblock-ai-dev-platform-public/new/main/Off-the-shelf-ready-made-dataset

Uses

<!-- Address questions around how the dataset is intended to be used. -->

Direct Use

<!-- This section describes suitable use cases for the dataset. -->

This dataset is released for the sole purpose of fine-tuning AI models. Commercial use and resale are strictly prohibited.

Out-of-Scope Use

<!-- This section addresses misuse, malicious use, and uses that the dataset will not work well for. --> Commercial use and resale are strictly prohibited.

Dataset Structure

<!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. --> Each conversation contains 2 separated audios, stereo, no overlap. Topic name is mentioned as the folder name.

Dataset Creation

Curation Rationale

<!-- Motivation for the creation of this dataset. -->

This dataset was created as a premium service for us to train a ASR system.

Source Data

<!-- This section describes the source data (e.g. news text and headlines, social media posts, translated sentences, ...). -->

Data Collection and Processing

<!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. -->

Native USA speakers in different age groups, gender recorded stereo audios, spoke spontaneously.

Who are the source data producers?

<!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. -->

AIxBlock

Recommendations

<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->

Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.

Dataset Card Authors [optional]

AIxBlock

Dataset Card Contact

Join our Discord to ask for more info: https://discord.gg/nePjg9g5v6 Or create a request on our github here: https://github.com/AIxBlock-2023/aixblock-ai-dev-platform-public/tree/main