datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vox2-face-landmark
Dataset Card for "voxceleb2-features"
More Information needed
google-landmarkshow2sign-asl-landmarks
How2Sign ASL Landmarks
Sentence-level MediaPipe landmark cache for the How2Sign ASL translation dataset. This dataset stores preprocessed landmark arrays only, not MP4 clips.
The rows come from the local How2Sign realigned CSV files and front-view raw videos. Each successful row is one sentence segment sampled at a 25 FPS cap and stored in compressed NPZ shards.
Why This Dataset Exists
The public How2Sign data has known alignment/completeness issues: some text rows do not… See the full description on the dataset page: https://huggingface.co/datasets/martinctl/how2sign-asl-landmarks.google_landmarks_places
Google Landmarks places
Google Landmarks is a great dataset, but it lacks geospatial information about the places. This dataset fills
this gap by providing latitude and longitude for each landmark. The dataset also contains the name of the landmark from OpenStreetMap
and information about the country, the province/state, and the city/village where the landmark is located. This information was collected from OSM
via Nominatim.
landmarks
Dataset Card for "landmarks"
More Information needed
Optimized_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this Parquet file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Optimized_Video_Facial_Landmarks.juno-landmark-rftgoogle_landmarks_photos
Dataset Card for "google_landmarks_photos"
More Information needed
slovo-full-landmarks
Slovo landmarks base
Base landmarks dataset built from full Slovo videos with MediaPipe Tasks HolisticLandmarker.
google_landmark_v2_qaEmotion_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this CSV file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Emotion_Video_Facial_Landmarks.CASL-W60-Landmarks
CASL-W60-64-frames Landmarks
license: mit
task_categories:
- video-classification
language:
- en
tags:
- sign-language
- mediapipe
- skeleton
- Central-African-Sign-Language
pretty_name: CASL-W60 Landmarks
size_categories:
- 1K<n<10K
Dataset Card for CASL-W60 Landmarks
This dataset contains preprocessed holistic landmarks for Central African Sign Language (CASL). It is specifically designed for training sequence-based models like Transformers or LSTMs for Sign… See the full description on the dataset page: https://huggingface.co/datasets/luciayen/CASL-W60-Landmarks.how2sign-landmarks-front-raw-parquet
Split
Count
train
31047
validation
1739
test
2343
Parquet-shared front mediapipe data from PSewmuthu/How2Sign_Holistic
import cv2
import numpy as np
import tempfile
import os
from datasets import load_dataset
from IPython.display import Video, display
# 1. Initialize Streams
print("Initializing streams...")
landmark_stream = load_dataset("bdanko/how2sign-landmarks-front-raw-parquet", split="train", streaming=True)
rgb_stream =… See the full description on the dataset page: https://huggingface.co/datasets/bdanko/how2sign-landmarks-front-raw-parquet.landmark-en-hed
Dataset Card for "landmark-en-hed"
More Information needed
welsh-speech-landmarks
Welsh Speech Dataset - Facial Landmarks
68-point facial landmarks (ibug68 template) from the Welsh Speech Dataset.
Contents
Facial landmarks for every frame
68 3D points per frame (x, y, z coordinates)
Format: Parquet
Manual annotation using ibug68 template
Format
The landmarks.parquet file contains:
Column
Description
speaker_id
Speaker identifier (1-33)
phrase_id
Phrase identifier (1-10)
frame_id
Frame identifier (e.g., "001", "002")… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-landmarks.Taichi-HD-Landmarks
A subset of Taichi-HD dataset with estimated body landmarks, ported from https://github.com/AliaksandrSiarohin/motion-cosegmentation.
The protocol of dataset creation can be found at the original paper: https://arxiv.org/abs/2004.03234
We backup the raw images (./raw_assets) and landmarks (./landmark) in this repo.
Landmark_And_FER2013Emotion_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this CSV file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/mac26/Emotion_Video_Facial_Landmarks.bio_landmarks_zh_gemma-2-9b-itwlasl-upper-arm-pose-landmarks
WLASL Pose Landmarks (MediaPipe Lite)
This dataset contains 3D pose landmarks extracted from the WLASL (World Level American Sign Language) video dataset using Google MediaPipe. It is designed to facilitate lightweight sign language recognition models by providing pre-computed skeletal data instead of raw video pixels.
Supported Tasks
Sign Language Recognition (SLR)
Skeleton-based Action Recognition
Keypoint Analysis
Dataset Structure
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Kibalama/wlasl-upper-arm-pose-landmarks.HandGesture_LandmarkCoordinates_Skeleton
HandGesture_LandmarkCoordinates_Skeleton
Just a datasets for my essay
Very simple with:
Coordinates from 21 Landmark (x0-x20, y0-y20)
8 different labels (thumb_up, thumb_down, point_up, ok_sign, peace, open_hand, nothing, other)
Change Log:
29/08/2025: Uploaded little datasets
17/09/2025: Updated from 2449 rows to 18629 rows
muse-landmark-1500ChicagoFSWild-Landmarksaahq-face-landmarks-30kaahq-face-landmarks-full-meshentity-visual-landmark_all_Qwen2.5-VL-7B-Instructlandmark-m3bio_landmarks_zh_gemma-2-9b-it_testnyc-landmark-descriptions
Dataset Card for NYC Landmark Descriptions
This dataset card documents the NYC Landmark Descriptions dataset.It contains 100+ manually written landmark descriptions, each ~200 characters long, labeled with vibe and a binary touristy tag. The augmented split expands to 1,000 samples.
Dataset Details
Dataset Description
Curated by: Bareethul Kader (Carnegie Mellon University)
Language(s): English
License: CC BY 4.0
Repository: bareethul/nyc-landmark-descriptions… See the full description on the dataset page: https://huggingface.co/datasets/bareethul/nyc-landmark-descriptions.how2sign-landmarks-curated-normalizedA specially selected set of 96 mediapipe coordinates from Mediapipe Holistic.
import numpy as np
from datasets import load_dataset
dataset_name = "bdanko/how2sign-landmarks-curated-normalized"
print(f"Loading dataset {dataset_name}...")
dataset = load_dataset(dataset_name, split="test", streaming=True)
# Get a sample
sample = next(iter(dataset))
video_id = sample['video_id']
sentence = sample['sentence']
features_raw = sample['features']
shape = sample['shape']
# Decode features
landmarks… See the full description on the dataset page: https://huggingface.co/datasets/bdanko/how2sign-landmarks-curated-normalized.
