datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shona1realtime-turn-detection-test-data
Realtime speech test recordings
Synthetic speech recordings for black-box Realtime API behavior tests in
Speaches. Each WAV file is the unmodified output of OpenAI
text-to-speech. Tests are responsible for adding silence, combining recordings, and choosing streaming chunk
boundaries for their scenarios.
metadata.jsonl follows the Hugging Face AudioFolder layout. Each record contains the generation inputs, file
digest, expected text, transcription, and word/speech intervals from… See the full description on the dataset page: https://huggingface.co/datasets/speaches-ai/realtime-turn-detection-test-data.pine-realtime-1.0-preview-tau2-voice-trajectories
Pine Realtime 1.0 Preview — tau2-bench voice trajectories, audio & recordings
Full trajectories, benchmark-side audio, and agent-side recordings for
pine-realtime-1.0-preview evaluated on all 278 tau2-bench voice tasks
(50 airline, 114 retail, 114 telecom).
Results (Pass@1)
Domain
Pass@1
airline
78.0%
retail
78.1%
telecom
78.1%
mean
78.0%
Simulated user: openrouter/openai/gpt-5.2; user TTS: ElevenLabs eleven_v3; 1200s per-task budget.… See the full description on the dataset page: https://huggingface.co/datasets/pine-ai/pine-realtime-1.0-preview-tau2-voice-trajectories.real-time-voice
Real-Time Voice AI Hears but Does Not Listen
Audio recordings from the paper Real-Time Voice AI Hears but Does Not Listen.
Code: https://github.com/bartelds/real-time-voice
Subsets
multiturn — one .wav per run: a single recording of the full conversation, with the
caller audio and the model's spoken responses concatenated in turn order. Columns: model, scenario,
instruction, condition, run index.
single_turn — diagnostic stimulus clips. Columns: label… See the full description on the dataset page: https://huggingface.co/datasets/bartelds/real-time-voice.shona2
