CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kyutai /Babillage Babillage Babillage is a multimodal benchmark dataset introduced along with MoshiVis (Project Page | arXiv), containing three common vision-language benchmarks converted in spoken form, for the evaluation of Vision Speech Models. For each benchmark (COCO-Captions, OCR-VQA, VQAv2), we first reformat the text question-answer pairs into a more conversational dialogue, and then convert them using a text-to-speech pipeline, using a consistent synthetic voice for the answer (assistant)… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/Babillage.audiovisual-question-answering100K<n<1M14 likes407 downloads2y agoHugging Face02kyutai /Audio-NTREX-4L Audio-NTREX-4L Dataset Description Audio-NTREX-4L is a long-form multilingual speech translation dataset from 🇫🇷 French, 🇪🇸 Spanish, 🇵🇹 Portuguese and 🇩🇪 German to 🇬🇧 English designed to evaluate speech translation models on multi-sentence utterances. It is built from the text translation dataset NTREX by aggregating multiple sentences from a same context to create new source texts and their reference translation. We then use 3 different state-of-the-art… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/Audio-NTREX-4L.audiotranslation1K<n<10K6 likes327 downloads8mo agoHugging Face03kyutai /voices_tts_longeval Voices TTS-Longeval A set a voices for kyutai-labs/tts_longeval. Voice samples are taken from: libri/: the LibriSpeech ASR Corpus, released under CC BY 4.0 license. seed/: taken from ButedanceSpeech/seed-tts-eval, the voices seem to come from CommonVoice, released under CC 0 license. ntrex_dialogs/en: taken from VCTK, released Creative Commons, Attribution 4.0 International. ntrex_dialogs/fr: taken from the CML-TTS dataset, released under the Creative Commons License:… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/voices_tts_longeval.audiotext-to-speech1K<n<10K1 likes178 downloads1y agoHugging Face04kyutai /hifitts2-aligned HiFiTTS-2 word alignments Word-level forced alignments for the HiFiTTS-2 corpus (44 kHz subset, resampled to 24 kHz), as used to train pocket-tts models. Like HiFiTTS-2 itself, this dataset contains no audio — only pointers and annotations. The audio is downloaded from LibriVox and cut locally. Contents train/train_aligned-*.jsonl.gz — the full aligned training manifest eval_aligned.jsonl.gz — a 1000-utterance held-out split scripts/download_audio.py — fetches… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/hifitts2-aligned.tabulartext-to-speech10M<n<100M2 likes154 downloads1mo agoHugging Face05kyutai /HaluEvalAudio_1000 HaluEvalAudio 1000 Dataset Dataset Description HaluEvalAudio 1000 is a specialized speech-based question-answering dataset designed to benchmark the capabilities of general multimodal & audio-focused language models as well as retrieval-augmented audio language models. Compared to common QA benchmarks such as Llama Questions, Web Questions, or TriviaQA, HaluEvalAudio 1000 introduces more challenging questions and topics and is specifically structured for… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/HaluEvalAudio_1000.audioaudio-text-to-text1K<n<10K1 likes88 downloads5mo agoHugging Face06kyutai /KairosQA KairosQA Dataset Dataset Description KairosQA is a temporally grounded question-answering dataset designed to evaluate the temporal alignment and reasoning capabilities of Large Language Models (LLMs). Unlike static benchmarks, KairosQA focuses on facts that evolve over time, specifically subject–relation–object triplets from Wikidata that changed at least twice between 2018 and 2025 as described in our paper Understanding Data Temporality Impact on Large Language… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/KairosQA.textquestion-answering1K<n<10K1 likes77 downloads4mo agoHugging Face07Atufa /codeswitch-fr-en-kyutai-stt Code-Switched French–English STT Probe Dataset Dataset Summary This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.audioautomatic-speech-recognitionn<1K0 likes60 downloads7mo agoHugging Face08kalbin /moshi-on-policy-dpo-v20-kyutai-alignedaudio1K<n<10K0 likes32 downloads5mo agoHugging Face09open-llm-leaderboard /kyutai__helium-1-preview-2b-detailsgated Dataset Card for Evaluation run of kyutai/helium-1-preview-2b Dataset automatically created during the evaluation run of model kyutai/helium-1-preview-2b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kyutai__helium-1-preview-2b-details.tabular10K<n<100K0 likes30 downloads2y agoHugging Face10kalbin /moshi-on-policy-dpo-v20-kyutai-smokeaudion<1K0 likes11 downloads5mo agoHugging Face11kalbin /moshi-on-policy-dpo-v20-kyutai-aligned-smokeaudion<1K0 likes4 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.