CoolFace
Datasetpublic

Appenlimited/1000h-us-english-smartphone-conversation

πŸ“š 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings) This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for: Automatic Speech Recognition (ASR) Speaker Identification and Gender/Age Analysis Dialect and Accent Modeling Multi-speaker Speech Separation 🧾 Dataset Contents The dataset includes: metadata.CSV: Metadata including speaker… See the full description on the dataset page: https://huggingface.co/datasets/Appenlimited/1000h-us-english-smartphone-conversation.

sourceHugging Faceunknownupdated 1y agoView on Hugging Face
3likes135downloads
Dataset Card

πŸ“š 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings)

This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for:

  • β€”Automatic Speech Recognition (ASR)
  • β€”Speaker Identification and Gender/Age Analysis
  • β€”Dialect and Accent Modeling
  • β€”Multi-speaker Speech Separation

🧾 Dataset Contents

The dataset includes:

  • β€”metadata.CSV: Metadata including speaker gender, age, nationality, etc.
  • β€”TRANSCRIPTION_AUTO_SEGMENTED: Automatically segmented transcriptions
  • β€”COPYRIGHT.TXT / README.TXT: Copyright notice and original description
  • β€”Transcription_Conventions.pdf: Transcription and annotation guidelines

πŸ’‘ Use Cases

  • β€”Teaching / Demonstrating Speech Annotation
  • β€”Research in Speech Analysis
  • β€”Training or Fine-tuning Small ASR Models

⚠️ Usage Notes

This dataset was collected by Appen. For copyright details, please refer to COPYRIGHT.TXT. Unauthorized use for commercial purposes is prohibited.

πŸ§‘β€πŸ’» Citation Recommendation

If you use this dataset in a paper or project, please cite it as:

"USE-ASR003 Dataset Sample, Appen Butler Hill Pty Ltd, 2018."

Let me know if you need any adjustments or a more formal version!