Appenlimited/1000h-us-english-smartphone-conversation
π 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings) This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for: Automatic Speech Recognition (ASR) Speaker Identification and Gender/Age Analysis Dialect and Accent Modeling Multi-speaker Speech Separation π§Ύ Dataset Contents The dataset includes: metadata.CSV: Metadata including speakerβ¦ See the full description on the dataset page: https://huggingface.co/datasets/Appenlimited/1000h-us-english-smartphone-conversation.
π 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings)
This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for:
- Automatic Speech Recognition (ASR)
- Speaker Identification and Gender/Age Analysis
- Dialect and Accent Modeling
- Multi-speaker Speech Separation
π§Ύ Dataset Contents
The dataset includes:
- metadata.CSV: Metadata including speaker gender, age, nationality, etc.
- TRANSCRIPTION_AUTO_SEGMENTED: Automatically segmented transcriptions
- COPYRIGHT.TXT / README.TXT: Copyright notice and original description
- Transcription_Conventions.pdf: Transcription and annotation guidelines
π‘ Use Cases
- Teaching / Demonstrating Speech Annotation
- Research in Speech Analysis
- Training or Fine-tuning Small ASR Models
β οΈ Usage Notes
This dataset was collected by Appen. For copyright details, please refer to COPYRIGHT.TXT. Unauthorized use for commercial purposes is prohibited.
π§βπ» Citation Recommendation
If you use this dataset in a paper or project, please cite it as:
"USE-ASR003 Dataset Sample, Appen Butler Hill Pty Ltd, 2018."
Let me know if you need any adjustments or a more formal version!
