CoolFace
Datasetpublic

Codyfederer/bbbb

bbbb This is a merged speech dataset containing 345 audio segments from 2 source datasets. Dataset Information Total Segments: 345 Speakers: 7 Languages: en Emotions: happy, sad, neutral, angry Original Datasets: 2 Dataset Structure Each example contains: audio: Audio file (WAV format, 16kHz sampling rate) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged datasets) emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/bbbb.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes14downloads
Dataset Card

bbbb

This is a merged speech dataset containing 345 audio segments from 2 source datasets.

Dataset Information

  • —Total Segments: 345
  • —Speakers: 7
  • —Languages: en
  • —Emotions: happy, sad, neutral, angry
  • —Original Datasets: 2

Dataset Structure

Each example contains:

  • —audio: Audio file (WAV format, 16kHz sampling rate)
  • —text: Transcription of the audio
  • —speaker_id: Unique speaker identifier (made unique across all merged datasets)
  • —emotion: Detected emotion (neutral, happy, sad, etc.)
  • —language: Language code (en, es, fr, etc.)

Usage

Loading the Dataset

python
from datasets import load_dataset

# Load the dataset
dataset = load_dataset("Codyfederer/bbbb")

# Access the training split
train_data = dataset["train"]

# Example: Get first sample
sample = train_data[0]
print(f"Text: {sample['text']}")
print(f"Speaker: {sample['speaker_id']}")
print(f"Language: {sample['language']}")
print(f"Emotion: {sample['emotion']}")

# Play audio (requires audio libraries)
# sample['audio']['array'] contains the audio data
# sample['audio']['sampling_rate'] contains the sampling rate

Alternative: Load from CSV

python
import pandas as pd
from datasets import Dataset, Audio, Features, Value

# Load the CSV file
df = pd.read_csv("data.csv")

# Define features
features = Features({
    "audio": Audio(sampling_rate=16000),
    "text": Value("string"),
    "speaker_id": Value("string"),
    "emotion": Value("string"),
    "language": Value("string")
})

# Create dataset
dataset = Dataset.from_pandas(df, features=features)

Dataset Structure

The dataset includes:

  • —data.csv - Main dataset file with all columns
  • —segments/ - Directory containing all audio files
  • —load_dataset.txt - Python script for loading the dataset (rename to .py to use)

CSV columns:

  • —audio: Path to the audio file (in segments/ directory)
  • —text: Transcription of the audio
  • —speaker_id: Unique speaker identifier
  • —emotion: Detected emotion
  • —language: Language code

Speaker ID Mapping

Speaker IDs have been made unique across all merged datasets to avoid conflicts. For example:

  • —Original Dataset A: speaker_0, speaker_1
  • —Original Dataset B: speaker_0, speaker_1
  • —Merged Dataset: speaker_0, speaker_1, speaker_2, speaker_3

Original dataset information is preserved in the metadata for reference.

Data Quality

This dataset was created using the Vyvo Dataset Builder with:

  • —Automatic transcription and diarization
  • —Quality filtering for audio segments
  • —Music and noise filtering
  • —Emotion detection
  • —Language identification

License

This dataset is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0).

Citation

bibtex
@dataset{vyvo_merged_dataset,
  title={bbbb},
  author={Vyvo Dataset Builder},
  year={2025},
  url={https://huggingface.co/datasets/Codyfederer/bbbb}
}

This dataset was created using the Vyvo Dataset Builder tool.