beatsprom/audio-speech-realtime-voice-agents-2026
ποΈ Audio, Speech Foundation Models & Real-Time Voice Agents Dataset (2026 Edition) A structured research dataset featuring 1,722 domain-verified research papers and 298 official code repositories focused on Full-Duplex Speech-to-Speech LLMs, Real-Time Voice Agents (<200ms Latency), Zero-Shot TTS, Voice Cloning, OpenAI Whisper-v3, Neural Audio Codecs (EnCodec/DAC/SNAC), and Generative Music (2023β2026). Built with Universal Scientific Engine V18.1 Diamond, providing 48 schemaβ¦ See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/audio-speech-realtime-voice-agents-2026.
ποΈ Audio, Speech Foundation Models & Real-Time Voice Agents Dataset (2026 Edition)
A structured research dataset featuring 1,722 domain-verified research papers and 298 official code repositories focused on Full-Duplex Speech-to-Speech LLMs, Real-Time Voice Agents (<200ms Latency), Zero-Shot TTS, Voice Cloning, OpenAI Whisper-v3, Neural Audio Codecs (EnCodec/DAC/SNAC), and Generative Music (2023β2026).
Built with Universal Scientific Engine V18.1 Diamond, providing 48 schema attributes with verified repository attribution, 8 AI topological semantic clusters, pre-calculated Top-3 Semantic Nearest Neighbors Graph, structured benchmark leaderboards, and native 384-dimensional dense PyTorch embeddings.
π Dataset Schema Highlights (48 Columns)
π§© 8 AI Semantic Clusters Breakdown
Neural Audio Codecs & Discrete Acoustic Tokenizers(375 papers)End-to-End Multilingual Speech Recognition (ASR & Whisper)(254 papers)Speech Enhancement, Denoising & Source Separation(247 papers)Audio-Visual Speech Synthesis, Talking Heads & Lip-Sync(233 papers)Generative Music, Audio Diffusion & Sound Synthesis(198 papers)Low-Latency Zero-Shot TTS & Instant Voice Cloning(191 papers)Spoken Dialogue Evaluation & Acoustic Benchmarking(137 papers)Full-Duplex Voice Agents & Speech-to-Speech LLMs(87 papers)
π Interactive OpenAngels Visual Dashboard Included
Open DATASET_ANALYTICS_DASHBOARD_100_SAMPLE.html directly in your browser (Chrome/Edge/Safari) to explore the interactive visual intelligence directory with real-time filtering, search, and audio latency metrics.
π» 1-Click Python Quickstart
import pyarrow.parquet as pq
# Load 100-Sample Teaser
table = pq.read_table("AUDIO_SPEECH_FOUNDATION_MODELS_REALTIME_VOICE_AGENTS_2026_100_SAMPLE.parquet")
df = table.to_pandas()
print(f"Loaded {len(df)} sample Audio & Speech AI papers.")
print(f"Top Paper: {df['title'].iloc[0]}")
print(f"Task Paradigm: {df['audio_speech_task_paradigm'].iloc[0]}")
print(f"Latency Profile: {df['speech_latency_profile'].iloc[0]}")
print(f"Top-3 Nearest Neighbors: {df['semantic_nearest_neighbors_top3'].iloc[0]}")π Get the Full 1,722-Paper Enterprise Edition
The complete commercial production dataset (1,722 papers in Parquet with 384d vectors, SQLite DB, Clean CSV, Interactive OpenAngels HTML Dashboard, and JSON) is available here:
π [BeatsProm Audio, Speech & Real-Time Voice Agents Dataset Full Edition](https://beatsprom.gumroad.com/l/audio-speech-real-time-voice-agents-2026)
