datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DailyTalkContiguous
DailyTalkContiguous
This repo contains a concatenated version of the DailyTalk dataset (official repo).
Rather than having separate files for each speaker's turn, this uses a stereo file for each conversation. The two speakers in a conversation
are put separately on the left and right channels.
The dataset is annotated with word level timestamps.
The original DailyTalk dataset and baseline code are freely available for academic use with CC-BY-SA 4.0 license, this dataset
uses the… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/DailyTalkContiguous.Babillage
Babillage
Babillage is a multimodal benchmark dataset introduced along with MoshiVis (Project Page | arXiv), containing three common vision-language benchmarks converted in spoken form, for the evaluation of Vision Speech Models.
For each benchmark (COCO-Captions, OCR-VQA, VQAv2), we first reformat the text question-answer pairs into a more conversational dialogue, and then convert them using a text-to-speech pipeline, using a
consistent synthetic voice for the answer (assistant)… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/Babillage.interactivity-alignment-samples
Audio Samples: Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models
Audio samples accompanying the paper "Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models".
Paper: arxiv.org
Blog post: kyutai.org
Models: 🤗 huggingface.co
Overview
This repository hosts the audio samples generated on Full-Duplex-Bench v1 (static evaluation with pre-recorded input) and Full-Duplex-Bench v2 (real-time multi-turn dialogue with GPT-Realtime), used in… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/interactivity-alignment-samples.Audio-NTREX-4L
Audio-NTREX-4L
Dataset Description
Audio-NTREX-4L is a long-form multilingual speech translation dataset from 🇫🇷 French, 🇪🇸 Spanish, 🇵🇹 Portuguese and 🇩🇪 German to 🇬🇧 English designed to evaluate speech translation models on multi-sentence utterances. It is built from the text translation dataset NTREX by aggregating multiple sentences from a same context to create new source texts and their reference translation. We then use 3 different state-of-the-art… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/Audio-NTREX-4L.voices_tts_longeval
Voices TTS-Longeval
A set a voices for kyutai-labs/tts_longeval.
Voice samples are taken from:
libri/: the LibriSpeech ASR Corpus, released under CC BY 4.0 license.
seed/: taken from ButedanceSpeech/seed-tts-eval,
the voices seem to come from CommonVoice, released under CC 0 license.
ntrex_dialogs/en: taken from VCTK, released
Creative Commons, Attribution 4.0 International.
ntrex_dialogs/fr: taken from the CML-TTS dataset, released under the
Creative Commons License:… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/voices_tts_longeval.librispeech_test_clean_enhancedHaluEvalAudio_1000
HaluEvalAudio 1000 Dataset
Dataset Description
HaluEvalAudio 1000 is a specialized speech-based question-answering dataset designed to benchmark the capabilities of general multimodal & audio-focused language models as well as retrieval-augmented audio language models.
Compared to common QA benchmarks such as Llama Questions, Web Questions, or TriviaQA, HaluEvalAudio 1000 introduces more challenging questions and topics and is specifically structured for… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/HaluEvalAudio_1000.codeswitch-fr-en-kyutai-stt
Code-Switched French–English STT Probe Dataset
Dataset Summary
This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.moshi-on-policy-dpo-v20-kyutai-alignedmoshi-on-policy-dpo-v20-kyutai-smokemoshi-on-policy-dpo-v20-kyutai-aligned-smokekyutai-tts-voices
