datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio-Transcription-Models-Comparison-PT-BR
Audio Transcription Models Comparison
A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese.
About the Dataset
This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering:
Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.Eng-Filipino-Accented-audio-with-human-transcription-call-center-topicThis dataset contains 103+ hours of spontaneous English conversations spoken in a Filipino accent, recorded in a studio environment to ensure crystal-clear audio quality. The conversations are designed as role-play scenarios between agents and customers across a variety of call center domains.
🗣️ Speech Style: Natural, unscripted role-playing between native Filipino-accented English speakers, simulating real-world customer interactions.
🎧 Audio Format: High-quality stereo WAV files, recorded… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Eng-Filipino-Accented-audio-with-human-transcription-call-center-topic.Thai-H2H-Call-center-audio-with-human-transcriptionThis dataset contains natural Thai-language conversations between human agents and human customers, designed to reflect realistic call center interactions across multiple domains. All conversations are conducted through unscripted role-playing, allowing for spontaneous and dynamic exchanges that closely mirror real-world scenarios.
🗣️ Speech Type: Human-to-human dialogues simulating customer-agent interactions.
🎭 Style: Non-scripted, spontaneous role-playing to capture authentic speech… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Thai-H2H-Call-center-audio-with-human-transcription.English-USA-NY-Boston-AAVE-Audio-with-transcriptionThis dataset captures spontaneous English conversations from native U.S. speakers across distinct regional and cultural accents, including:
🗽 New York English
🎓 Boston English
🎤 African American Vernacular English (AAVE)
The recordings span three real-life scenarios:
General Conversations – informal, everyday discussions between peers.
Call Center Simulations – customer-agent style interactions mimicking real support environments.
Media Dialogue – scripted reads and semi-spontaneous… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/English-USA-NY-Boston-AAVE-Audio-with-transcription.massive-audio-transcription-pipeline
massive-audio-transcription-pipeline outputs
Transcription outputs from the
massive-audio-transcription-pipeline,
a parallel Whisper pipeline that chunks long audio into overlapping windows,
transcribes across a worker pool, merges lightweight speaker diarization, and
checkpoints every chunk for crash resume.
Generation method
Backend: faster-whisper base model (CTranslate2), 1 worker.
Audio: real public-domain speech from the Hugging Face LibriSpeech dummy… See the full description on the dataset page: https://huggingface.co/datasets/narinzar/massive-audio-transcription-pipeline.
