CoolFace
Datasetpublic

ud-nlp/british-english-speech-recognition-dataset

British English Telephone Dialogues Dataset - 200 Hours The dataset consists of 200 hours of high-quality telephone dialogues from 310 native speakers in the UK, with detailed annotations (transcriptions, timestamps, speaker ID, gender, and background noise) to support speech recognition systems, NLP tasks, and machine learning models requiring diverse British English audio datasets. - Get the data Dataset characteristics: Characteristic Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/british-english-speech-recognition-dataset.

sourceHugging Facecc-by-nc-nd-4.0updated 8mo agoView on Hugging Face
0likes30downloads
Dataset Card

British English Telephone Dialogues Dataset - 200 Hours

The dataset consists of 200 hours of high-quality telephone dialogues from 310 native speakers in the UK, with detailed annotations (transcriptions, timestamps, speaker ID, gender, and background noise) to support speech recognition systems, NLP tasks, and machine learning models requiring diverse British English audio datasets. - [Get the data](https://unidata.pro/datasets/british-english-speech-recognition-dataset/?utm_source=huggingface-nlp&utm_medium=referral&utm_campaign=british-english-speech-recognition-dataset)

Dataset characteristics:

CharacteristicData
DescriptionAudio of telephone dialogues in English for training NLP models in real-world conversational scenarios.
Data typesAudio
TasksSpeech recognition, NLP
CountryThe United Kingdom (GBK)
Hours of telephone dialogue200
Number of speakers310
LabelingAnnotation (text content, speaker's ID, gender, age and other attributes)
GenderMale (42%), Female (58%)
Recording deviceAndroid smartphone, iPhone

📊 Sample dataset available! For full access, contact us to discuss purchase terms.

Dataset structure

  • —audio - audio file
  • —text - text transcription
  • —British English Speech Recognition Dataset.csv - metadata for the data

🧩 Like the dataset but need different data? We can collect a custom dataset just for you - learn more about our data collection services here

Similar Datasets:

  1. 1.American Speech Recognition Dataset
  2. 2.LLM Text Generation Dataset
  3. 3.Spanish Speech Recognition Dataset

🌐 UniData - your trusted data partner. Unique, accurate, thoroughly collected and annotated data designed to fuel your AI/ML success.