CoolFace
20 results

kyutai

kyutai /DailyTalkContiguous DailyTalkContiguous This repo contains a concatenated version of the DailyTalk dataset (official repo). Rather than having separate files for each speaker's turn, this uses a stereo file for each conversation. The two speakers in a conversation are put separately on the left and right channels. The dataset is annotated with word level timestamps. The original DailyTalk dataset and baseline code are freely available for academic use with CC-BY-SA 4.0 license, this dataset uses the… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/DailyTalkContiguous.audio21 likes2.4k downloads2y agoHugging Facekyutai /rocket-sciencegated Rocket Science Time-aligned video, keyboard actions, game events, and per-frame game state for all four players, captured from a 2v2 Rocket League match. This is the dataset behind MIRA, a real-time multiplayer world model trained to simulate Rocket League gameplay — by General Intuition and Kyutai, in collaboration with Epic Games. Code, technical report, and a live demo: github.com/mira-wm/mira · mira-wm.com. Loading from mira.data import RocketScienceDataset… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/rocket-science.videoother1M<n<10M58 likes1.7k downloads2mo agoHugging Facekyutai /Babillage Babillage Babillage is a multimodal benchmark dataset introduced along with MoshiVis (Project Page | arXiv), containing three common vision-language benchmarks converted in spoken form, for the evaluation of Vision Speech Models. For each benchmark (COCO-Captions, OCR-VQA, VQAv2), we first reformat the text question-answer pairs into a more conversational dialogue, and then convert them using a text-to-speech pipeline, using a consistent synthetic voice for the answer (assistant)… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/Babillage.audiovisual-question-answering100K<n<1M14 likes407 downloads2y agoHugging Facekyutai /interactivity-alignment-samples Audio Samples: Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models Audio samples accompanying the paper "Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models". Paper: arxiv.org Blog post: kyutai.org Models: 🤗 huggingface.co Overview This repository hosts the audio samples generated on Full-Duplex-Bench v1 (static evaluation with pre-recorded input) and Full-Duplex-Bench v2 (real-time multi-turn dialogue with GPT-Realtime), used in… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/interactivity-alignment-samples.audio1K<n<10K9 likes404 downloads4mo agoHugging Facekyutai /librispeech-enhanced-voice-prompts LibriSpeech enhanced voice prompts The voice-cloning prompts used to evaluate pocket-tts / Kyutai TTS models on the LibriSpeech-PC test-clean cross-sentence protocol. 448 unique prompt utterances from LibriSpeech test-clean, speech-enhanced (denoised, 32 kHz mono FLAC, duration-preserving), laid out as <speaker>/<chapter>/<utterance>.flac — the same tree structure as LibriSpeech itself, so --prompt-root style substitution works directly. Also included:… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/librispeech-enhanced-voice-prompts.text-to-speech0 likes336 downloads1mo agoHugging Facekyutai /Audio-NTREX-4L Audio-NTREX-4L Dataset Description Audio-NTREX-4L is a long-form multilingual speech translation dataset from 🇫🇷 French, 🇪🇸 Spanish, 🇵🇹 Portuguese and 🇩🇪 German to 🇬🇧 English designed to evaluate speech translation models on multi-sentence utterances. It is built from the text translation dataset NTREX by aggregating multiple sentences from a same context to create new source texts and their reference translation. We then use 3 different state-of-the-art… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/Audio-NTREX-4L.audiotranslation1K<n<10K6 likes327 downloads7mo agoHugging Face