datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arkitscenes_mcmc_3dgs
Data Statistics
Scenes
Mean PSNR ↑
Mean SSIM ↑
Mean LPIPS ↓
Mean Depth L1 ↓
Mean #3DGS
Total #3DGS
1,290
31.63 dB
0.907
0.217
0.0051 m
1.149M
1.483B
audio2face-mediapipe-arkit-teacher
audio2face-mediapipe-arkit-teacher
Left: source video frame (face-cropped). Middle: MediaPipe FaceLandmarker's 478 landmark points. Right: an illustrative subset of mp_bs — the 52-channel ARKit blendshape vector shipped in this dataset — as horizontal bars updating per frame.
14,703 emotional-speech clips, each annotated with a 52-channel ARKit blendshape sequence extracted by MediaPipe FaceLandmarker from the source video (or from audio-driven synthesis where no… See the full description on the dataset page: https://huggingface.co/datasets/myned-ai/audio2face-mediapipe-arkit-teacher.audio2face-emotion-arkit-teacher
audio2face-emotion-arkit-teacher
Nyx avatar (Gaussian-splat head, ARKit-52 blendshape rig) driven by a surprise clip's blendshape labels derived from this dataset.
14,082 emotional-speech clips, each annotated with two parallel 52-channel ARKit blendshape sequences (NVIDIA Audio2Face-3D-v2.3.1-James and LAM_Audio2Expression) plus a 26-dimensional NVIDIA Audio2Emotion conditioning vector.
Reference-only dataset — the original audio is not shipped. Each row contains a… See the full description on the dataset page: https://huggingface.co/datasets/myned-ai/audio2face-emotion-arkit-teacher.arkitscenes-preprocesseditems_raw_fullitems_raw_lite
