full-duplex
otoSpeech-full-duplex-turn-104h
Dataset Card for otoSpeech-full-duplex-turn-104h
Contact
Website: https://oto.earthEmail: agent@oto.earth
Dataset Summary
otoSpeech-full-duplex-turn-104h is an English, full-duplex conversational speech dataset for research on turn-taking and related spoken-dialogue phenomena. It contains 420 two-speaker conversations totaling approximately 104.94 hours. Each conversation includes time-aligned, channel-separated audio, a stereo combined recording… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-turn-104h.otoSpeech-full-duplex-280h
📢 Notice: We released a new processed version of otoSpeech. Available here: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-processed-141h
Dataset Card for otoSpeech-full-duplex-280h: Full-Duplex Conversational Speech Dataset
Contact
Website: https://oto.earth
Mail: consome@oto.earth
Dataset Summary
otoSpeech-full-duplex-280h is a 280-hour, full-duplex, two-speaker conversational speech dataset.
Each sample includes 48 kHz… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-280h.video-full-duplex-benchmark
VideoFDB: Video-Full-Duplex-Benchmark
Project Page · HuggingFace · Paper (arXiv)
Dataset Description
A benchmark dataset of annotated, two-person video conference recordings designed to support the evaluation of multimodal AI agents in conversational settings. The dataset covers 11 distinct conversational dynamics — spanning verbal, nonverbal, and mixed-modality behavior — annotated through a three-pass human-in-the-loop pipeline.
The benchmark consists of trimmed… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/video-full-duplex-benchmark.otoSpeech-full-duplex-processed-141h
Dataset Card for otoSpeech-full-duplex-processed-141h: Full-Duplex Conversational Speech Dataset
Dataset Summary
otoSpeech-full-duplex-processed-141h is a full-duplex, two-speaker conversational speech dataset. It is derived from otoSpeech-full-duplex-280h and has been curated and processed as follows:
Selected high-quality conversations based on human reviews.
Applied noise reduction and speech enhancement to improve audio quality.
Added new samples collected after the… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-processed-141h.Full-Duplex-Bench-DataOmni-DuplexEval-Full
Omni-DuplexEval
Omni-DuplexEval is a benchmark for evaluating real-time duplex multimodal interaction. Unlike conventional offline video understanding benchmarks, Omni-DuplexEval focuses on streaming settings where models must continuously process evolving multimodal inputs and decide what to respond and when to respond.
The benchmark contains two scenarios:
Real-Time Description (RTD): evaluates continuous streaming description ability.
Proactive Reminder (PR): evaluates… See the full description on the dataset page: https://huggingface.co/datasets/tianrantianran/Omni-DuplexEval-Full.
