datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chatbot-arena-spoken-voicesspeech-chatbot-alpaca-eval
This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
python audio_evals/main.py --dataset speech-chatbot-alpaca-eval --model gpt4o_speech
🚀超凡体验,尽在UltraEval-Audio🚀
UltraEval-Audio——全球首个同时支持语音理解和语音生成评估的开源框架,专为语音大模型评估打造,集合了34项权威Benchmark,覆盖语音、声音、医疗及音乐四大领域,支持十种语言,涵盖十二类任务。选择UltraEval-Audio,您将体验到前所未有的便捷与高效:
一键式基准管理 📥:告别繁琐的手动下载与数据处理,UltraEval-Audio为您自动化完成这一切,轻松获取所需基准测试数据。
内置评估利器… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/speech-chatbot-alpaca-eval.chatbotarena-spoken-all-7824
ChatbotArena-Spoken Dataset
Based on ChatbotArena, we employ GPT-4o-mini to select dialogue turns that are well-suited to spoken-conversation analysis, yielding 7824 data points. To obtain audio, we synthesize every utterance, user prompt and model responses, using one of 12 voices from KokoroTTS (v0.19), chosen uniformly at random. Because the original human annotations assess only lexical content in text, we keep labels unchanged and treat them as ground truth for the spoken… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/chatbotarena-spoken-all-7824.chatbot-arena-spoken-voices
