datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Human-to-machine-Japanese-audio-call-center-conversations
Dataset Card for Japanese audio call center human to machine conversations
This dataset contains synthetic audio conversations in Japanese between human customers and machine agents, simulating real-world call center scenarios
Dataset Details
Dataset Description
Curated by: AIxBlock (aixblock.io)
Funded by [optional]: AIxBlock (aixblock.io)
Shared by [optional]: AIxBlock (aixblock.io)
Language(s) (NLP): Japanese
License: Creative Commons Attribution Non… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Human-to-machine-Japanese-audio-call-center-conversations.Gemma-4-E4B_hidden_to_audio_tokens-2.0
Gemma-4 S2S Alignment Dataset (Multilingual) – Version 2.0 (260K)
This dataset is specifically engineered to train a lightweight, low-latency Hidden-to-Speech (H2S) alignment model.
By capturing the raw, abstract semantic representations from the 33rd hidden states of a Text LLM (Gemma-4 8B) and mapping them directly onto quantized discrete audio streams, this dataset bypasses traditional text generation bottlenecks to establish native Speech-to-Speech (S2S) processing… See the full description on the dataset page: https://huggingface.co/datasets/firdavsus/Gemma-4-E4B_hidden_to_audio_tokens-2.0.audio_to_blendshapes_maintibetan-to-english-audio-dataset
Tibetan to English Audio Dataset
Dataset Description
A Tibetan speech recognition dataset with transcriptions and English translations containing 1,642 audio samples.
Dataset Summary
This dataset contains Tibetan speech recordings with:
Tibetan transcriptions in native script
English translations
High-quality audio files in WAV format
Total Samples: 1,178Total Size: ~1.1 GBAudio Format: WAV
Languages
Source Language: Tibetan (བོད་སྐད་)
Target… See the full description on the dataset page: https://huggingface.co/datasets/Titung/tibetan-to-english-audio-dataset.audio-to-midi-test-kit
Audio to MIDI Test Kit
Synthetic audio fixtures, observed MIDI output and a review checklist for checking browser-based audio-to-MIDI conversion. This is a small reproducibility kit, not an accuracy benchmark or a comparison of competing products.
The observations were collected using the browser tool at Note From Audio.
Contents
fixtures/: one original six-second melody encoded as MP3, PCM WAV, AAC-in-M4A, FLAC, Vorbis-in-OGG and ADTS AAC.
results/: the actual… See the full description on the dataset page: https://huggingface.co/datasets/yihoukeji/audio-to-midi-test-kit.tibetan-audio-to-english-fixed-filtered
Tibetan audio translation Dataset
Dataset Description
Tibetan audio translation Dataset
Dataset Summary
This dataset contains 6,366 audio samples with corresponding transcriptions, totaling approximately 15.8 hours of audio.
Languages
The dataset is in EN (Language code: en).
Dataset Structure
Data Fields
audio: An audio object containing:
path: Path to the audio file (if applicable)
array: Audio waveform as a numpy array… See the full description on the dataset page: https://huggingface.co/datasets/Titung/tibetan-audio-to-english-fixed-filtered.audio-to-blendshapes-longer-datasetaudio_to_blendshapes_testGemma-4-E4B_hidden_to_audio_tokens
Gemma-4 S2S Alignment Dataset (Multilingual)
This dataset is specifically engineered to train a lightweight, low-latency Hidden-to-Speech (H2S) alignment model.
By capturing the raw, abstract semantic representations from the last hidden states of a Text LLM (Gemma-4 8B) and mapping them directly onto quantized discrete audio streams, we can bypass traditional text generation bottlenecks to establish native Speech-to-Speech (S2S) processing pipelines.
Key Conceptual… See the full description on the dataset page: https://huggingface.co/datasets/firdavsus/Gemma-4-E4B_hidden_to_audio_tokens.audio-text-embed-to-imagesSpectrogram_Audio_text_to_Base64Thai-human-to-machine-call-center-audio-with-scriptThis dataset features natural Thai-language conversations between human speakers and machine agents, simulating real-world call center interactions across a variety of customer service domains. All dialogues are non-scripted and performed as role-play scenarios, capturing spontaneous, realistic exchanges.
🗣️ Speech Type: Human-to-machine conversations, simulating AI agents and human customers' dialogues.
🎭 Style: Spontaneous, unscripted role-playing, designed to reflect actual customer… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Thai-human-to-machine-call-center-audio-with-script.Bass_Audio_text_to_Base64audio-to-text-datasetaudio-to-image-sample-knowledge-base
Sample Knowledge Base Dataset
This folder contains a small starter dataset for the audio-to-image retrieval project.
It has 15 educational diagram images and metadata that can be used to build a Pinecone vector index.
Files
data/sample_knowledge_base/
train/
images/
*.png
metadata.csv
README.md
The images are generated educational diagrams for machine learning topics.
What Each Image Record Needs
Each image should have these fields:… See the full description on the dataset page: https://huggingface.co/datasets/Anuragleo67/audio-to-image-sample-knowledge-base.sanskrit_audio_dataset_under_30_taged_meta_to_text_from_edgesample_text_to_audio_audiosample_text_to_audio_audio
