datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
realtime-conversational-voice-agent-duplex-2026
🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026)
This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab.
The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.vi-fdb-v1-gpt-realtime
GPT-Realtime on Vi-FDB v1: full reference outputs
Public reference outputs from GPT-Realtime on all 400 Vi-FDB event cases and
their 200 available clean controls. This repository contains model outputs
and evaluation evidence, not benchmark inputs.
Canonical benchmark: https://huggingface.co/datasets/tuanamz/vi-fdb-v1
Benchmark revision: c9730eddf78b094610be09796afe8e60c178caef
Interactive explorer: https://huggingface.co/spaces/tuanamz/vi-fdb-v1-explorer
Contents… See the full description on the dataset page: https://huggingface.co/datasets/tuanamz/vi-fdb-v1-gpt-realtime.audio-speech-realtime-voice-agents-2026
🎙️ Audio, Speech Foundation Models & Real-Time Voice Agents Dataset (2026 Edition)
A structured research dataset featuring 1,722 domain-verified research papers and 298 official code repositories focused on Full-Duplex Speech-to-Speech LLMs, Real-Time Voice Agents (<200ms Latency), Zero-Shot TTS, Voice Cloning, OpenAI Whisper-v3, Neural Audio Codecs (EnCodec/DAC/SNAC), and Generative Music (2023–2026).
Built with Universal Scientific Engine V18.1 Diamond, providing 48 schema… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/audio-speech-realtime-voice-agents-2026.
