CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rorosese /my-voxtral-datasetaudion<1K0 likes87 downloads1y agoHugging Face02MrlolDev /voxtral-emotion-speech Voxtral Emotion Speech Dataset Emotional speech dataset generated with ElevenLabs v3 using audio tags for emotion control, validated with SenseVoice for quality assurance. Quality Filter Each generated clip is validated using SenseVoice (iic/SenseVoiceSmall) to ensure the emotion in the audio matches the expected label. Process: Generate audio with ElevenLabs v3 using emotion audio tags Run SenseVoice inference on the audio Compare detected emotion with expected emotion… See the full description on the dataset page: https://huggingface.co/datasets/MrlolDev/voxtral-emotion-speech.audioaudio-classification1K<n<10K0 likes47 downloads5mo agoHugging Face03trishtan /voxtral-forensic-dsaudio1K<n<10K1 likes29 downloads7mo agoHugging Face04MrlolDev /voxtral-emotion-temporal VoxTral Emotion Temporal Dataset Dataset for training emotion transition detection in speech. ~500 clips with frame-level emotion annotations at 20ms resolution. Overview Property Value Clips ~500 Sample Rate 16000 Hz Frame Resolution 20ms Emotions 6 (neutral, happy, angry, sad, surprise, fear) Clip Distribution Single Emotion (40%): One emotion throughout Transitions (40%): Emotion changes mid-sentence (2-3 segments) No Emotion (20%):… See the full description on the dataset page: https://huggingface.co/datasets/MrlolDev/voxtral-emotion-temporal.audion<1K0 likes29 downloads5mo agoHugging Face05shakods /voxtral-synthetic-eng-test Voxtral Synthetic English (ASR) Synthetic speech dataset for fine-tuning Voxtral ASR models. English utterances generated with ElevenLabs TTS from the CohereLabs/aya_collection_language_split (english, targets column). All audio is 16 kHz mono WAV. Dataset structure Column Type Description audio_path string Path to the audio file in this repo (e.g. audio/utt_000000.wav) text string Ground-truth transcript for the audio Audio: 16 kHz, mono, WAV, stored… See the full description on the dataset page: https://huggingface.co/datasets/shakods/voxtral-synthetic-eng-test.audioautomatic-speech-recognitionn<1K0 likes27 downloads7mo agoHugging Face06Tonic /voxtral-dataset-20250913_174653 Voxtral ASR Dataset This dataset was created using the Voxtral ASR Fine-tuning Interface. Dataset Structure audio_path: Relative path to the audio file (stored in audio/ directory) text: Transcription of the audio Dataset Statistics Number of examples: 10 Audio files uploaded: 10 Total dataset size: 10,015,616 bytes Usage from datasets import load_dataset, Audio # Load dataset dataset = load_dataset("Tonic/voxtral-dataset-20250913_174653") #… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/voxtral-dataset-20250913_174653.audion<1K0 likes23 downloads1y agoHugging Face07coyotte508 /voxtral-tts-french-samplesSamples from https://huggingface.co/mistralai/Voxtral-4B-TTS-2603 audion<1K0 likes22 downloads6mo agoHugging Face08Tonic /voxtral-dataset-20250913_175651 Voxtral ASR Dataset This dataset was created using the Voxtral ASR Fine-tuning Interface. Dataset Structure audio_path: Relative path to the audio file (stored in audio/ directory) text: Transcription of the audio Dataset Statistics Number of examples: 10 Audio files uploaded: 10 Total dataset size: 10,015,583 bytes Usage from datasets import load_dataset, Audio # Load dataset dataset = load_dataset("Tonic/voxtral-dataset-20250913_175651") #… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/voxtral-dataset-20250913_175651.audion<1K0 likes13 downloads1y agoHugging Face09Trelis /eval-Voxtral-Mini-3B-2507-medical-terms-2025-20260408-1926 Evaluation Results: Voxtral-Mini-3B-2507 Evaluation results from Whisper model evaluation. Summary Model WER CER mistralai/Voxtral-Mini-3B-2507 8.14% 3.37% Source Data Evaluation Dataset: Trelis/medical-terms-2025 Model Evaluated: mistralai/Voxtral-Mini-3B-2507 Columns Column Description audio Audio sample (if available from source dataset) reference Ground truth transcription prediction Model prediction wer Word… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-Voxtral-Mini-3B-2507-medical-terms-2025-20260408-1926.audion<1K0 likes12 downloads6mo agoHugging Face10trishtan /voxtral-forensic-ds-splitstext1K<n<10K1 likes11 downloads7mo agoHugging Face11Trelis /eval-Voxtral-Mini-3B-2507-eka-hard-20260408-1920 Evaluation Results: Voxtral-Mini-3B-2507 Evaluation results from Whisper model evaluation. Summary Model WER CER mistralai/Voxtral-Mini-3B-2507 43.86% 29.54% Source Data Evaluation Dataset: Trelis/eka-hard Model Evaluated: mistralai/Voxtral-Mini-3B-2507 Columns Column Description audio Audio sample (if available from source dataset) reference Ground truth transcription prediction Model prediction wer Word Error Rate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-Voxtral-Mini-3B-2507-eka-hard-20260408-1920.audion<1K0 likes10 downloads6mo agoHugging Face12Tonic /voxtral-dataset-20260225_222829 Voxtral ASR Dataset This dataset was created using the Voxtral ASR Fine-tuning Interface. Dataset Structure audio_path: Relative path to the audio file (stored in audio/ directory) text: Transcription of the audio Dataset Statistics Number of examples: 10 Audio files uploaded: 10 Total dataset size: 13,328,128 bytes Usage from datasets import load_dataset, Audio # Load dataset dataset = load_dataset("Tonic/voxtral-dataset-20260225_222829") #… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/voxtral-dataset-20260225_222829.audion<1K0 likes9 downloads7mo agoHugging Face13Trelis /eval-Voxtral-Mini-3B-2507-multimed-hard-20260408-1931 Evaluation Results: Voxtral-Mini-3B-2507 Evaluation results from Whisper model evaluation. Summary Model WER CER mistralai/Voxtral-Mini-3B-2507 10.89% 7.46% Source Data Evaluation Dataset: Trelis/multimed-hard Model Evaluated: mistralai/Voxtral-Mini-3B-2507 Columns Column Description audio Audio sample (if available from source dataset) reference Ground truth transcription prediction Model prediction wer Word Error… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-Voxtral-Mini-3B-2507-multimed-hard-20260408-1931.audion<1K0 likes7 downloads6mo agoHugging Face14Rcarvalo /voxtral-french-1000h0 likes2 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.