datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Real-TurnTurk
Real-TurnTurk
English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the… See the full description on the dataset page: https://huggingface.co/datasets/tugrulbayrak/Real-TurnTurk.LibriReplay-DOA
LibriReplay-DOA (Anonymous Submission)
Overview
LibriReplay-DOA is a multi-channel multi-speaker replay dataset designed for evaluating
robust speech processing systems under realistic room playback conditions.
The dataset contains replay recordings captured in real rooms under
multiple playback configurations (DOA settings). Each session includes
multiple overlapping speakers.
This dataset is released for peer-review purposes.
Dataset Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/real-recordings/LibriReplay-DOA.
