datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fdb-v1-outputs-v1
Full-Duplex-Bench v1.0 model outputs (fdb-v1-outputs-v1)
7997 model responses = 11 benchmark runs × the 727 stimuli of
Full-Duplex-Bench v1.0 (pause handling,
backchannel, smooth turn-taking, user interruption). The runs cover 7 systems;
several differ only in voice prompt, prompting regime or weights, which is the point — those are
controlled pairs. For every stimulus and run you get the
model's own reply channel as lossless FLAC — time-synchronous with the stimulus, so… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/fdb-v1-outputs-v1.ifbench-conversations-v1
IF-Bench conversations (ifbench-conversations-v1)
3000 synthetic full-duplex spoken conversations: 15 examiner configurations
(9 speech models), each holding the same 200-task set of Full-Duplex-Bench v2 staged scenarios as the
examiner (the model under study — it carries a role, a topic and four goals to hit in order)
against a PersonaPlex-7B examinee that is never told the topic. Per dialogue you get both
channels as lossless mono FLAC, the exact prompts/voices/sampling… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/ifbench-conversations-v1.
