datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
realtime-conversational-voice-agent-duplex-2026
🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026)
This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab.
The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.solana-clawd-realtime-research-instruct
Solana Clawd Realtime Research Instruct
Instruction-tuning dataset generated by scripts/realtime_dataset_ingest.py
from submitted PDFs, notebooks, parquet QA rows, JSON/JSONL files, and local
reference text.
Contents
Total examples: 29058
Train/eval/test: 26152 / 1452 / 1454
Sources: 28
Duplicate examples removed: 0
Duplicate files skipped: 2
Secret-like records skipped: 296
Format
Each row uses OpenAI/Hugging Face chat messages:
{"messages":… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-realtime-research-instruct.fastapi_websockets_realtime_backpressure_teaser
🚀 Python Backend - FastAPI WebSockets & Real-Time Connection Backpressure Triage (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Python Backend - FastAPI WebSockets & Real-Time Connection Backpressure Triage on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
500 Verified FAANG… See the full description on the dataset page: https://huggingface.co/datasets/emgena/fastapi_websockets_realtime_backpressure_teaser.autoinference-realtime-mix-v1
Autoinference Real-Time Generation Mix v1
This is a prompt set for the real_time_generation serving benchmark. That profile
stands in for medium-context, single-shot interactive traffic: roughly 3000 input
tokens, 100 output tokens, one request at a time with no shared context between
requests. The usual way to run it feeds the server random token IDs of a fixed
length. This dataset keeps the same input and output shape but uses real prompts.
The reason real text matters: random… See the full description on the dataset page: https://huggingface.co/datasets/modal-labs/autoinference-realtime-mix-v1.
