datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audio-3dvgHowFarAreYou_3DSpeakerTrain_fullAudio2Face-3D-Dataset-v1.0.0-claire
Dataset Description:
NVIDIA Audio2Face-3D-dataset-v1.0.0-claire includes audio files, blendshape data, animated geometry caches, geometry files, and transform files.
This dataset is for demonstration purposes and not for production usage.
For source code, documentation, helper scripts, packaged builds, and links to all components in the Audio2Face-3D technology stack, visit the Audio2Face-3D GitHub repository
Dataset Owner:
NVIDIA Corporation
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Audio2Face-3D-Dataset-v1.0.0-claire.combined_3datasets_801010HowFarAreYou_3DSpeakerTrain_full
Dataset Card for "HowFarAreYou_3DSpeakerTrain"
More Information needed
vox2_3D_distill_shard22vox2_3D_distill_shard24vox2_3D_distill_shard21vox2_3D_distill_shard23vox2_3D_distill_shard13vox2_3D_distill_shard17vox2_3D_distill_shard16vox2_3D_distill_shard40vox2_3D_distill_shard42vox2_3D_distill_shard14vox2_3D_distill_shard32vox2_3D_distill_shard20HowFarAreYou_3DSpeakerTrain
Dataset Card for "HowFarAreYou_3DSpeakerTrain"
More Information needed
vox2_3D_combined_shard_2vox2_3D_distill_shard_06HowFarAreYou_3DSpeakerTrain
Dataset Card for "HowFarAreYou_3DSpeakerTrain"
More Information needed
vox2_3D_shard00343D-CAVFA
3D-CAVFA Dataset
Overview
The 3D-CAVFA dataset is a multimodal collection featuring synchronized audio and facial blendshape coefficients captured from 20 subjects, totaling 15 hours. Its linguistic content encompasses frequently used modern Chinese words and phrases, alongside a variety of sentence patterns drawn from daily life, thereby guaranteeing strong relevance to real-world scenarios.
Citation
If you use this dataset, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/TianshunHan/3D-CAVFA.vox2_3D_distill_shard_18Emergence-Text-Image-Audio-3D
Emergence: The Four Forms of Intelligence
Summary
A multimodal dataset that unifies Text, Image, Audio, and 3D modalities with quad-modality alignment for every sample, ensuring that each record contains semantically consistent representations of the same concept.
This dataset is curated by using 3D assets from Objaverse as anchors and aligning them with semantically corresponding images and audio clips from various sources using a embedding search… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Emergence-Text-Image-Audio-3D.HowFarAreYou_3DSpeaker
Dataset Card for "HowFarAreYou_3DSpeaker"
More Information needed
MEAD_3D_shard_06vox2_3D_distill_shard01MEAD_3D_W035mac-m4pro-confirmation-3drive-20260903
sbpn-mac-m4pro-confirmation-three-drive-20260902
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/mac-m4pro-confirmation-3drive-20260903.
