CoolFace
20 results

audio-to-audio

AIxBlock /Human-to-machine-Japanese-audio-call-center-conversations Dataset Card for Japanese audio call center human to machine conversations This dataset contains synthetic audio conversations in Japanese between human customers and machine agents, simulating real-world call center scenarios Dataset Details Dataset Description Curated by: AIxBlock (aixblock.io) Funded by [optional]: AIxBlock (aixblock.io) Shared by [optional]: AIxBlock (aixblock.io) Language(s) (NLP): Japanese License: Creative Commons Attribution Non… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Human-to-machine-Japanese-audio-call-center-conversations.audiotext-to-speechn<1K2 likes61 downloads1y agoHugging Facesanjuhs /audio_to_blendshapes_maintextn<1K0 likes43 downloads1y agoHugging Facefirdavsus /Gemma-4-E4B_hidden_to_audio_tokens-2.0 Gemma-4 S2S Alignment Dataset (Multilingual) – Version 2.0 (260K) This dataset is specifically engineered to train a lightweight, low-latency Hidden-to-Speech (H2S) alignment model. By capturing the raw, abstract semantic representations from the 33rd hidden states of a Text LLM (Gemma-4 8B) and mapping them directly onto quantized discrete audio streams, this dataset bypasses traditional text generation bottlenecks to establish native Speech-to-Speech (S2S) processing… See the full description on the dataset page: https://huggingface.co/datasets/firdavsus/Gemma-4-E4B_hidden_to_audio_tokens-2.0.text100K<n<1M0 likes43 downloads3mo agoHugging FaceTitung /tibetan-to-english-audio-dataset Tibetan to English Audio Dataset Dataset Description A Tibetan speech recognition dataset with transcriptions and English translations containing 1,642 audio samples. Dataset Summary This dataset contains Tibetan speech recordings with: Tibetan transcriptions in native script English translations High-quality audio files in WAV format Total Samples: 1,178Total Size: ~1.1 GBAudio Format: WAV Languages Source Language: Tibetan (བོད་སྐད་) Target… See the full description on the dataset page: https://huggingface.co/datasets/Titung/tibetan-to-english-audio-dataset.audioautomatic-speech-recognition1K<n<10K0 likes33 downloads8mo agoHugging Faceyihoukeji /audio-to-midi-test-kit Audio to MIDI Test Kit Synthetic audio fixtures, observed MIDI output and a review checklist for checking browser-based audio-to-MIDI conversion. This is a small reproducibility kit, not an accuracy benchmark or a comparison of competing products. The observations were collected using the browser tool at Note From Audio. Contents fixtures/: one original six-second melody encoded as MP3, PCM WAV, AAC-in-M4A, FLAC, Vorbis-in-OGG and ADTS AAC. results/: the actual… See the full description on the dataset page: https://huggingface.co/datasets/yihoukeji/audio-to-midi-test-kit.0 likes30 downloads11d agoHugging FaceTitung /tibetan-audio-to-english-fixed-filtered Tibetan audio translation Dataset Dataset Description Tibetan audio translation Dataset Dataset Summary This dataset contains 6,366 audio samples with corresponding transcriptions, totaling approximately 15.8 hours of audio. Languages The dataset is in EN (Language code: en). Dataset Structure Data Fields audio: An audio object containing: path: Path to the audio file (if applicable) array: Audio waveform as a numpy array… See the full description on the dataset page: https://huggingface.co/datasets/Titung/tibetan-audio-to-english-fixed-filtered.audioautomatic-speech-recognition1K<n<10K0 likes27 downloads8mo agoHugging Face