datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
otoSpeech-full-duplex-turn-104h
Dataset Card for otoSpeech-full-duplex-turn-104h
Contact
Website: https://oto.earthEmail: agent@oto.earth
Dataset Summary
otoSpeech-full-duplex-turn-104h is an English, full-duplex conversational speech dataset for research on turn-taking and related spoken-dialogue phenomena. It contains 420 two-speaker conversations totaling approximately 104.94 hours. Each conversation includes time-aligned, channel-separated audio, a stereo combined recording… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-turn-104h.otoSpeech-full-duplex-280h
📢 Notice: We released a new processed version of otoSpeech. Available here: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-processed-141h
Dataset Card for otoSpeech-full-duplex-280h: Full-Duplex Conversational Speech Dataset
Contact
Website: https://oto.earth
Mail: consome@oto.earth
Dataset Summary
otoSpeech-full-duplex-280h is a 280-hour, full-duplex, two-speaker conversational speech dataset.
Each sample includes 48 kHz… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-280h.video-full-duplex-benchmark
VideoFDB: Video-Full-Duplex-Benchmark
Project Page · HuggingFace · Paper (arXiv)
Dataset Description
A benchmark dataset of annotated, two-person video conference recordings designed to support the evaluation of multimodal AI agents in conversational settings. The dataset covers 11 distinct conversational dynamics — spanning verbal, nonverbal, and mixed-modality behavior — annotated through a three-pass human-in-the-loop pipeline.
The benchmark consists of trimmed… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/video-full-duplex-benchmark.otoSpeech-full-duplex-processed-141h
Dataset Card for otoSpeech-full-duplex-processed-141h: Full-Duplex Conversational Speech Dataset
Dataset Summary
otoSpeech-full-duplex-processed-141h is a full-duplex, two-speaker conversational speech dataset. It is derived from otoSpeech-full-duplex-280h and has been curated and processed as follows:
Selected high-quality conversations based on human reviews.
Applied noise reduction and speech enhancement to improve audio quality.
Added new samples collected after the… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-processed-141h.Full-Duplex-Bench-DataOmni-DuplexEval-Full
Omni-DuplexEval
Omni-DuplexEval is a benchmark for evaluating real-time duplex multimodal interaction. Unlike conventional offline video understanding benchmarks, Omni-DuplexEval focuses on streaming settings where models must continuously process evolving multimodal inputs and decide what to respond and when to respond.
The benchmark contains two scenarios:
Real-Time Description (RTD): evaluates continuous streaming description ability.
Proactive Reminder (PR): evaluates… See the full description on the dataset page: https://huggingface.co/datasets/tianrantianran/Omni-DuplexEval-Full.otoSpeech-full-duplex-task-oriented-20h
Dataset Viewer
https://cc-task-oriented-preview.vercel.app/
Task Walkthrough
https://www.oto.earth/research/task-oriented-dataset.html
What each of the seven tasks is for, what the two speakers could each see, and
how the interaction log lines up with the audio.
otoSpeech-full-duplex-task-oriented-20h
Contact
This sample dataset is provided for research purposes. We maintain larger and
more diverse datasets.
For collaborations… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-task-oriented-20h.korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.214-Hours-Korean-Full-Duplex-Multi-Channel-Speech-Dataset
Description
주제 기반 대화로 수집한 한국어(대한민국) 멀티스트림 자연 발화 스마트폰 음성 데이터셋입니다. 각 음성 데이터에는 대화 내용을 전사한 텍스트와 화자 ID, 성별, 연령 등의 속성 정보가 포함되어 있습니다. 본 데이터셋은 다양한 지역과 배경을 가진 폭넓은 화자들로부터 수집되었으며, 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1704?source=hf.kr
Specifications
Format
16 kHz, 16 bit, WAV, 모노 채널
Content category
정해진 주제 없이 자유롭게 진행된 자연 대화
Recording condition
낮은… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/214-Hours-Korean-Full-Duplex-Multi-Channel-Speech-Dataset.268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample
Description
Arabic(Saudi) Multi-stream Spontaneous Dialogue Smartphone speech dataset-Customer Service. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(268 native speakers), geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1627?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample.trusted-full-duplex-agent-data
Trusted Full-Duplex Speech Agent — Training Data & Evidence
Companion dataset for
model: jatshi/trusted-full-duplex-agent
and GitHub: Jatshi/trusted-full-duplex-agent.
2.0 evidence update
The 2.0 release adds the evidence needed to reproduce the deployed runtime rather
than only the offline research pipeline: real-time latency telemetry, unified
guardrail voice samples, the exact turn-taking MLP reports, TrustGate/ASR runtime
configuration, and integrity… See the full description on the dataset page: https://huggingface.co/datasets/jatshi/trusted-full-duplex-agent-data.172-hours-American-English-Full-Duplex-Multi-Channel-Speech-DatasetDescription
주제 기반 대화로 수집한 미국 영어 멀티스트림 자연 발화 스마트폰 음성 데이터셋입니다. 각 음성 데이터에는 대화 내용을 전사한 텍스트와 화자 ID, 성별, 연령 등의 속성 정보가 포함되어 있습니다. 본 데이터셋은 다양한 지역과 배경을 가진 폭넓은 화자들로부터 수집되었으며, 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상에 활용할 수 있습니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1770?fromCategoryPage=1?source=hf.kr
Specifications
Format
16 kHz, 16 bit, 비압축 WAV, 모노 채널, 화자별 채널 분리
Content category
주제 기반 자연 대화
Recording condition
낮은 배경 소음 환경(실내)
Recording device
Android 스마트폰, iPhone
Country… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/172-hours-American-English-Full-Duplex-Multi-Channel-Speech-Dataset.268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Dataset
Description
사우디아라비아 아랍어(Arabic-Saudi) 멀티스트림 자연 대화 스마트폰 고객 서비스 음성 데이터셋입니다. 다양한 고객 서비스 상황에서 자유롭게 대화하는 방식으로 수집되었으며, 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 총 268명의 아랍어 원어민 화자로부터 데이터를 수집하여 다양한 화자 특성을 반영했으며, 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1627?source=hf.kr
Specifications
Format
16kHz, 16 bit, WAV, 모노 채널
Content category
정해진 주제 없이 자유롭게 진행된 자연 대화… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Dataset.lalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset
Description
211시간 규모의 태국어(Thai-Thailand) 풀 듀플렉스 자연 대화 스마트폰 음성 데이터셋으로, 일반적인 주제를 기반으로 대화를 시뮬레이션하여 녹음했습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 총 654명의 태국어 원어민 화자로부터 데이터를 수집했으며, 다양한 화자와 지역적 특성을 반영하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1594?source=hf.kr
Format
16kHz, 16 bit, WAV, 모노 채널
Content category
일반적인 주제를 기반으로 대화를 시뮬레이션하여 녹음
Recording condition
낮은… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset.500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset
Description
일본어(Japan) 48kHz 풀 듀플렉스 자연 대화 스마트폰 음성 데이터셋으로, 주어진 주제를 바탕으로 자유롭게 대화하는 방식으로 수집되었습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 다양한 지역의 폭넓은 화자로부터 데이터를 수집하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다. 또한 다양한 AI 기업을 통해 데이터 품질을 검증했습니다.
데이터 수집, 저장 및 활용 전 과정에서 개인정보 보호 및 관련 법규를 엄격하게 준수하며, 사용자의 개인정보와 법적 권리를 보호합니다. 본 데이터셋은 GDPR, CCPA, PIPL을 준수합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1971?source=hf.kr
Specifications… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset.sample-fullduplex423-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Dataset
Description
필리핀 영어(English-Philippine) 멀티스트림 자연 대화 스마트폰 음성 데이터셋으로, 주어진 주제를 바탕으로 대화하는 방식으로 수집되었습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 다양한 지역의 폭넓은 화자로부터 데이터를 수집하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1771?source=hf.kr
Specifications
Format
16kHz, 16 bit, 비압축 WAV, 모노 채널, 화자별 채널 분리
Content category
주어진 주제를 기반으로 대화
Recording condition
낮은… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/423-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Dataset.211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample
Description
211 Hours - Thai(Thailand) Full-Duplex Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(654 native speakers), geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample.1100-Hours-Tagalog-Full-Duplex-Multi-Channel-Speech-Data-Sample
1100-Hours-Tagalog-Full-Duplex-Multi-Channel-Speech-Data-Sample
Description
Tagalog Multi-stream Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1100-Hours-Tagalog-Full-Duplex-Multi-Channel-Speech-Data-Sample.200-Hours-Korean-Full-Duplex-Multi-Channel-Speech-Data-Sample
Description
Korean(Korea) Multi-stream Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1704?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/200-Hours-Korean-Full-Duplex-Multi-Channel-Speech-Data-Sample.sca_full_duplex600-Hours-American-English-Full-Duplex-Multi-Channel-Speech-Data-Sample
Description
American English Multi-stream Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks.
For
more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1770?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/600-Hours-American-English-Full-Duplex-Multi-Channel-Speech-Data-Sample.352-Hours-Urdu-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample
Description
Urdu Full-Duplex Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/352-Hours-Urdu-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample.500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample
Description
Japanese(Japan) 48khz Full-Duplex Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample.600-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Data-Sample
Description
English(Philippine) Multi-stream Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1771?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/600-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Data-Sample.Omni-DuplexEval-Full
Omni-DuplexEval
Omni-DuplexEval is a benchmark for evaluating real-time duplex multimodal interaction. Unlike conventional offline video understanding benchmarks, Omni-DuplexEval focuses on streaming settings where models must continuously process evolving multimodal inputs and decide what to respond and when to respond.
The benchmark contains two scenarios:
Real-Time Description (RTD): evaluates continuous streaming description ability.
Proactive Reminder (PR): evaluates… See the full description on the dataset page: https://huggingface.co/datasets/foragi/Omni-DuplexEval-Full.otoSpeech-full-duplex-task-oriented-v2
Dataset Viewer
https://cc-task-oriented-preview.vercel.app/
otoSpeech-full-duplex-task-oriented-v2
Contact
This sample dataset is provided for research purposes. We maintain larger and
more diverse datasets.
For collaborations, inquiries, custom data collection, or joint research,
contact consome@oto.earth.
Dataset Summary
otoSpeech-full-duplex-task-oriented-v2 contains 17.376131 hours of English,
full-duplex, two-speaker… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-task-oriented-v2.otoSpeech-full-duplex-turn-104h-postprocessedotoSpeech-HQ-full-duplex-samples
Dataset Card for otoSpeech-HQ-full-duplex-samples: Full-Duplex Conversational Speech Dataset Samples
Dataset Summary
otoSpeech-HQ-full-duplex-samples is a curated collection of high-quality full-duplex conversational speech samples designed for commercial and production-oriented use.
This repository is derived from a private subset of otoSpeech and features carefully selected English two-speaker conversations with enhanced audio quality. The samples are intended for… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-HQ-full-duplex-samples.
