CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JavisVerse /MM-PreTrain JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation [HomePage] [Paper] [GitHub] TL;DR We introduce JavisGPT, a multimodal LLM that can understand audiovisual inputs and simultaneously generate synchronized sounding videos in a unified model. We also curate the JavisInst-Omni dataset to facilitate instruction-tuning for comprehension and generation on sounding videos. 📰 News [2025.12.30] 🚀 We release the training… See the full description on the dataset page: https://huggingface.co/datasets/JavisVerse/MM-PreTrain.audio100K<n<1M0 likes198 downloads9mo agoHugging Face02JavisVerse /JavisInst-Omni JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation [HomePage] [Paper] [GitHub] TL;DR We introduce JavisGPT, a multimodal LLM that can understand audiovisual inputs and simultaneously generate synchronized sounding videos in a unified model. We also curate the JavisInst-Omni dataset to facilitate instruction-tuning for comprehension and generation on sounding videos. 📰 News [2025.12.30] 🚀 We release the training… See the full description on the dataset page: https://huggingface.co/datasets/JavisVerse/JavisInst-Omni.audio10K<n<100K2 likes175 downloads9mo agoHugging Face03ctaguchi /SLR35_javaneseaudio100K<n<1M0 likes108 downloads24d agoHugging Face04javi22 /SpanishPodcastsaudioaudio-to-audio0 likes54 downloads8mo agoHugging Face05javi22 /high_quality_spanish_speechaudio10K<n<100K1 likes38 downloads8mo agoHugging Face06shunyalabs /javanese-speech-datasetaudio1K<n<10K0 likes34 downloads1y agoHugging Face07cheikh1499 /javaneseaudio1K<n<10K0 likes31 downloads6d agoHugging Face08JavisVerse /JavisData-Audioaudio100K<n<1M0 likes26 downloads1y agoHugging Face09javaabu /dhivehi-shaafiu-speechgatedDhivehi Shaafiu Speech is a single speaker Dhivehi speech dataset created by [Javaabu Pvt. Ltd.](https://javaabu.com). The dataset contains around 16.5 hrs of text read by professional Maldivian narrator Muhammadh Shaafiu. The text used for the recordings were text scrapped from various Maldivian news websites.audioautomatic-speech-recognition1K<n<10K2 likes22 downloads3y agoHugging Face10Serialtechlab /dhivehi-javaabu-speech-parquetgatedaudio1K<n<10K0 likes22 downloads9mo agoHugging Face11Speech-data /Javanese-Speech-Dataset 🎧 Javanese Speech Dataset The Javanese Speech Dataset is a structured and scalable speech audio dataset designed to provide high-quality audio data for training modern AI and machine learning models. It includes 85 hours of audio data across 585 files, delivered in MP3 and WAV formats, with a total size of 104 MB. This well-balanced audio dataset offers diverse and representative voice data, with 51% female and 49% male speakers, and an age range spanning from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Javanese-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes21 downloads6mo agoHugging Face12octava /fork-google-openslr-javaneseaudio1K<n<10K0 likes16 downloads2y agoHugging Face13javaabu /dhivehi-majlis-speechgatedDhivehi Majlis Speech is a Dhivehi speech dataset created from data annotated by [Javaabu Pvt. Ltd.](https://javaabu.com). The dataset contains around 10.5 hrs of speech collected from parliament sessions at The Peoples Majlis of Maldives (Maldivian Parliament) consisting of audio from different MPs from 6 different sessions.audioautomatic-speech-recognition1K<n<10K1 likes13 downloads3y agoHugging Face14Javid42 /azerbaijani-speech-datasetaudio100K<n<1M0 likes12 downloads10mo agoHugging Face15kaizee1 /openslr_java_sundaaudio0 likes9 downloads10mo agoHugging Face16javaabu /dhivehi-khadheeja-speechgatedDhivehi Khadheeja Speech is a single speaker Dhivehi speech dataset created by [Javaabu Pvt. Ltd.](https://javaabu.com). The dataset contains around 20 hrs of text read by professional Maldivian narrator Khadheeja Faaz. The text used for the recordings were text scrapped from various Maldivian news websites.audioautomatic-speech-recognition1K<n<10K0 likes8 downloads2y agoHugging Face17MG17-JetPunk /JavierMileiaudion<1K0 likes7 downloads3y agoHugging Face18keinelust /kusjankou-mikola-javar-z-kalinaju-kacjaryna-jagorava Kusjankou Mikola, Javar z kalinaju, Kacjaryna Jagorava Metadata Original Title (Cyrillic): Кусянкоў Мікола, Явар з калінаю, Кацярына Ягорава Transliterated Title: Kusjankou Mikola, Javar z kalinaju, Kacjaryna Jagorava Audio Files: 93 MP3 files Format: Belarusian audiobook Description This is a Belarusian audiobook dataset containing 93 audio tracks. License Please check the original source for licensing information. audion<1K0 likes7 downloads4mo agoHugging Face19javorski /corte.zipaudion<1K0 likes6 downloads3y agoHugging Face20javiimts /A0001_S003_0_G0001_G0002_continuations_es Dataset A0001_S003_0_G0001_G0002 para Ultravox (ES) Clips: 177 Formato: JSONL + WAV en audio/ Subido automáticamente con upload_to_hf.py. audion<1K0 likes5 downloads1y agoHugging Face21javiercubas /orpheus-es-datasetaudion<1K1 likes4 downloads2y agoHugging Face22Javzanpagam /cv-mn-workshopaudion<1K2 likes4 downloads4mo agoHugging Face23OwLim /javaneseTrain-sample-1000audio1K<n<10K0 likes3 downloads1y agoHugging Face24Javid42 /general_knowledge_geminigatedaudio10K<n<100K1 likes2 downloads7mo agoHugging Face25javaabu /scaling_dhivehi_sttgatedaudio100K<n<1M0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.