turn-taking
turntaking-pretraining-it-multilingual-3cturntaking-multilingual-llama8b-2aTurn_taking_prediction_SWBDsemantic-turn-takingjapanese-wav2vec2-base-turntaking-CSJparakeet-turntaking-stage1-libri-top4-ctc-spk-ds1parakeet-turntaking-stage1-mixed-libri-interact-top4-ctc-spkparakeet-turntaking-stage1-libri-top4-ctc-spk
dualturn-otospeech-turn-taking
OtoSpeech Turn-Taking
Official DualTurn release of the otospeech corpus, with per-frame turn-taking labels and
Mimi speech codec features. Each row is one full conversation. Frame rate 12.5 Hz (80 ms per frame).
Paper: DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
Training code: github.com/anyreachai/dualturn
Model checkpoint: anyreach-ai/dualturn-qwen2.5-mimi-0.5B
Splits
Split
Sessions
train
896
val
111
test
113… See the full description on the dataset page: https://huggingface.co/datasets/anyreach-ai/dualturn-otospeech-turn-taking.turn-taking-dataset
turn-taking-detection-dataset
scripts/notebooks for turn taking detection dataset
Frame Feature extraction
Currently, extracting feature with mmaciton: current method needs to be fix. It is difference from the OAD convention.current method: please refer thisOnly taking the central frame among non-overlapping video snippets. In OAD task, all the frames in from snippets are converted into the mean of features from them.
To-dos (legacy, now maintain @ notion)… See the full description on the dataset page: https://huggingface.co/datasets/anonseoul/turn-taking-dataset.dualturn-switchboard-turn-taking
Switchboard Turn-Taking
Official DualTurn release of the switchboard corpus, with per-frame turn-taking labels and
Mimi speech codec features. Each row is one full conversation. Frame rate 12.5 Hz (80 ms per frame).
Paper: DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
Training code: github.com/anyreachai/dualturn
Model checkpoint: anyreach-ai/dualturn-qwen2.5-mimi-0.5B
Splits
Split
Sessions
train
1986
val
295
test
138… See the full description on the dataset page: https://huggingface.co/datasets/anyreach-ai/dualturn-switchboard-turn-taking.semantic-turn-taking-benchmark
Semantic Turn-Taking Benchmark
A curated evaluation benchmark for semantic turn-taking models in voice AI. Given a conversation context, predict what action the AI agent should take: speak, listen, continue speaking, or continue listening.
Unlike acoustic-based approaches (VAD, silence detection), this benchmark tests whether a model can make turn-taking decisions from text/semantic content alone.
Action Classes
Action
Description
start_speaking
User… See the full description on the dataset page: https://huggingface.co/datasets/anyreach-ai/semantic-turn-taking-benchmark.avcocktail-turn-taking
AVCocktail Turn-Taking
Turn-taking event labels (speaker shifts and holds) for pairwise speaker
conversations, derived from the AVCocktail (from MCoRec challenge) audio-visual cocktail-party dataset.
Files are per-session, per-speaker-pair JSON label files, kept in the same directory layout as the source dataset:
train/
central-train_channelmaps.pkl # so that the model knows to map to the correct speaker in a dyadic
session_00/
events-spk_3-spk_4.json
session_01/… See the full description on the dataset page: https://huggingface.co/datasets/lggvu/avcocktail-turn-taking.candor-turntaking-annotations
CANDOR - Turn-Taking Annotations
Speech transcription and turn-taking annotation dataset built from the CANDOR corpus using NVIDIA Canary-Qwen2.5B ASR.
Dataset Description
This dataset contains 172,591 transcribed speech segments from the CANDOR conversational speech corpus (1,656 conversations). Each segment is a per-speaker utterance with Canary ASR transcript, designed for turn-taking prediction research.
Source
Audio corpus: CANDOR (English conversational… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/candor-turntaking-annotations.
