voice-agents
voiceagent-ser-open
VoiceAgent SER — redistributable corpus subset
Dataset Description
This dataset supports speaker-disjoint, multi-corpus speech emotion
recognition (SER) research — specifically the open_ser_v1 manifest used to
train and evaluate UNIMUSE, a model that fuses a frozen WavLM-Large encoder
with hand-crafted acoustic features via cross-attention for six-class
emotion classification (neutral, happy, sad, angry, fearful, disgust) across
German, Bangla, and Polish speech.… See the full description on the dataset page: https://huggingface.co/datasets/Pomoika24/voiceagent-ser-open.audio-speech-realtime-voice-agents-2026
🎙️ Audio, Speech Foundation Models & Real-Time Voice Agents Dataset (2026 Edition)
A structured research dataset featuring 1,722 domain-verified research papers and 298 official code repositories focused on Full-Duplex Speech-to-Speech LLMs, Real-Time Voice Agents (<200ms Latency), Zero-Shot TTS, Voice Cloning, OpenAI Whisper-v3, Neural Audio Codecs (EnCodec/DAC/SNAC), and Generative Music (2023–2026).
Built with Universal Scientific Engine V18.1 Diamond, providing 48 schema… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/audio-speech-realtime-voice-agents-2026.grounding_voice_agentsvoice-agent-sft-v1
Voice-Agent SFT v1
Intended use: Downstream SFT fine-tuning for phone AI voice agents. NOT for pretraining.
This dataset targets models that must handle natural conversation + tool calling + function
calls (CRM lookups, web search, RAG DB queries) in a voice-agent context.
Built as a downstream specialization dataset for the
Cubix-AI diffusion-research project,
specifically for fine-tuning the Qwen3.5-4B-Base masked-diffusion LLM produced in Milestone 1.
Published under HF org… See the full description on the dataset page: https://huggingface.co/datasets/CubixAI/voice-agent-sft-v1.
