datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audio2face-mediapipe-arkit-teacher
audio2face-mediapipe-arkit-teacher
Left: source video frame (face-cropped). Middle: MediaPipe FaceLandmarker's 478 landmark points. Right: an illustrative subset of mp_bs — the 52-channel ARKit blendshape vector shipped in this dataset — as horizontal bars updating per frame.
14,703 emotional-speech clips, each annotated with a 52-channel ARKit blendshape sequence extracted by MediaPipe FaceLandmarker from the source video (or from audio-driven synthesis where no… See the full description on the dataset page: https://huggingface.co/datasets/myned-ai/audio2face-mediapipe-arkit-teacher.arkitscenes-spatiallm
ARKitScenes-SpatialLM Dataset
ARkitScenes dataset preprocessed in SpatialLM format for oriented object bouding boxes detection with LLMs.
Overview
This dataset is derived from ARKitScenes 5,047 real-world indoor scenes captured using Apple's ARKit framework, preprocessed and formatted specifically for SpatialLM training.
Data Extraction
Point clouds and layouts are compressed in zip files. To extract the files, run the following script:
cd arkitscenes-spatiallm… See the full description on the dataset page: https://huggingface.co/datasets/ysmao/arkitscenes-spatiallm.audio2face-emotion-arkit-teacher
audio2face-emotion-arkit-teacher
Nyx avatar (Gaussian-splat head, ARKit-52 blendshape rig) driven by a surprise clip's blendshape labels derived from this dataset.
14,082 emotional-speech clips, each annotated with two parallel 52-channel ARKit blendshape sequences (NVIDIA Audio2Face-3D-v2.3.1-James and LAM_Audio2Expression) plus a 26-dimensional NVIDIA Audio2Emotion conditioning vector.
Reference-only dataset — the original audio is not shipped. Each row contains a… See the full description on the dataset page: https://huggingface.co/datasets/myned-ai/audio2face-emotion-arkit-teacher.arkitscenes-crossview-inpaint
Carlhahaha/arkitscenes-crossview-inpaint
Cross-view pairs generated from inpainting results (Topomap; meta root: /mnt/NAS/data/jz4725/topomap).
Splits
Uploaded splits: test, train, validation
Schema (columns)
Image_a: Original image of sample A (datasets.Image)
Image_b: Original image of sample B (datasets.Image)
Inpaint_b: Inpainted image of sample B (datasets.Image)
mask_b: Mask used for inpainting on B (datasets.Image)
point_b: Normalized centroid of B's… See the full description on the dataset page: https://huggingface.co/datasets/Carlhahaha/arkitscenes-crossview-inpaint.xx_beat_arkit_moshi_2025_07_20_30fps_attnarkitscenes-preprocessedsynthetic_accented_englishArkitscenes-Spatiallm
ARKitScenes-SpatialLM Dataset
ARkitScenes dataset preprocessed in SpatialLM format for oriented object bouding boxes detection with LLMs.
Overview
This dataset is derived from ARKitScenes 5,047 real-world indoor scenes captured using Apple's ARKit framework, preprocessed and formatted specifically for SpatialLM training.
Data Extraction
Point clouds and layouts are compressed in zip files. To extract the files, run the following script:
cd arkitscenes-spatiallm… See the full description on the dataset page: https://huggingface.co/datasets/Gen3DF/Arkitscenes-Spatiallm.items_prompts_fullitems_raw_fullxx_beat_arkit_encodec2_2025_07_31_30fps_attnitems_raw_litexx_cbh_arkit_face_v1_moshi_2025_08_10_30fps_attnGrupoRTX
