datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turn-detection-vietnameseNguồn dữ liệu: vi-wiki-conversational-search
Tỷ lệ Complete:Incomplete = 244304:366451
Đã lưu 610755 samples vào training_data.csv
Đã lưu 6189 samples vào test_data.csv
turn-end-detection
Turn-end detection from real ASR prefixes, with audio
100,348 labelled end-of-turn decision points over
53,140 synthesized customer-service utterances, each one paired
with the 16 kHz audio it was cut from, plus the endpointing decisions
6 commercial endpointer configurations
made on the same audio.
The question each row poses is the one a voice agent has to answer continuously:
given everything heard so far, has the caller finished speaking? Ending the
turn too early talks over… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/turn-end-detection.realtime-turn-detection-test-data
Realtime speech test recordings
Synthetic speech recordings for black-box Realtime API behavior tests in
Speaches. Each WAV file is the unmodified output of OpenAI
text-to-speech. Tests are responsible for adding silence, combining recordings, and choosing streaming chunk
boundaries for their scenarios.
metadata.jsonl follows the Hugging Face AudioFolder layout. Each record contains the generation inputs, file
digest, expected text, transcription, and word/speech intervals from… See the full description on the dataset page: https://huggingface.co/datasets/speaches-ai/realtime-turn-detection-test-data.urdu-turn-detection-audio-v2
🗣️ Urdu Turn Detection (Audio Dataset V2)
This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech.
It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection.
🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-turn-detection-audio-v2.Urdu-Turn-Detection-10k
Urdu Turn Detection Dataset 🗣️
A high-quality dataset of 10,000 Urdu sentences labeled for Turn Detection (End-of-Turn). This dataset is designed to help conversational AI systems determine if a user has finished speaking (Complete) or is pausing/trailing off (Incomplete).
Dataset Details
Total Samples: 10,000
Language: Urdu (ur) - Nastaliq/Arabic Script only.
Cleanliness: - 100% Urdu Script (No Roman/English).
Avg. Sentence Length: - ~7.7 words (33 characters)… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/Urdu-Turn-Detection-10k.turn_detection_3k_zh这是一个专为儿童机器人聊天场景设计的语料数据集,由 Gemini 2.5 Pro 生成。数据集旨在充分模拟儿童的聊天行为,覆盖了广泛的主题和对话类型,包括单轮和三轮对话。
主要涵盖的场景类别包括:
学科与知识探索:语文、数学、英语、科学、历史、地理及常识问答。
文化与传统:民俗节日、神话传说、礼仪习惯。
兴趣爱好与玩乐:玩具、各类游戏(电子、棋类、户外)、收藏。
创意表达与想象:绘画手工、音乐歌舞、故事创作、幻想世界。
个人与社交生活:关于自己、家庭亲人、朋友同学、学校生活(非学术方面)。
日常生活与环境观察:天气季节、饮食食物、动植物观察、交通工具、周围环境事件。
媒体与娱乐:动画片、漫画绘本、电影、儿童歌曲故事音频、适龄网络内容。
对机器人的互动与探索:询问机器人基本信息、能力功能、情感互动、测试挑战及音量调节等。
涵盖以下场景
A. 学科与知识探索 (Academic & Knowledge Exploration)
语文 (Chinese Language Arts):
认字、写字、组词、造句
古诗词(背诵、含义、诗人故事)
成语(含义、故事、接龙)… See the full description on the dataset page: https://huggingface.co/datasets/justpluso/turn_detection_3k_zh.customer-turn-detection-rate-limit-safe-with-artifactsturn-detection-labeled-v3
🇷🇺 Russian Real-Estate Turn Detection (Probability Balanced)
This dataset is designed for training probability-based turn detection models for Russian conversational AI, specifically in the real-estate domain (renting, buying, inquiries).
It follows a contrastive approach: for every complete user query, there is a corresponding incomplete version. This forces the model to learn the subtle semantic and grammatical cues that signal the end of a turn versus a mid-sentence pause.… See the full description on the dataset page: https://huggingface.co/datasets/RAS1981/turn-detection-labeled-v3.urdu-turn-detection-audio-v2
🗣️ Urdu Turn Detection (Audio Dataset V2)
This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech.
It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection.
🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/urdu-turn-detection-audio-v2.turn-detection-labeled-v1
Russian Spoken Turn Detection Dataset
This dataset is designed for training End-of-Utterance (EOU) / Turn Detection models for Russian voice agents. It contains pairs of spoken utterances labeled as either complete (user finished speaking) or incomplete (user pausing mid-sentence).
It was generated to fine-tune Large Language Models (like Qwen, Llama, etc.) to recognize when to interrupt a user versus when to wait, specifically handling Russian hesitation markers ("эээ", "ну"… See the full description on the dataset page: https://huggingface.co/datasets/RAS1981/turn-detection-labeled-v1.Egyptian_Turn_Detection_Datasetcustomer-turn-detection-structured-with-artifactscustomer-turn-detection-rate-limit-safeturn-detection-jess-5kturn-detection-labeled-v2customer-turn-detection-1kcustomer-turn-detection-structuredcustomer-turn-detection-balanced-with-artifactsnamo-turn-detection-samplearabic-turn-detection-pilotcustomer-turn-detection-fixedcustomer-turn-detection-testcustomer-turn-detection-test-with-artifacts
