datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
welsh-speech-landmarks
Welsh Speech Dataset - Facial Landmarks
68-point facial landmarks (ibug68 template) from the Welsh Speech Dataset.
Contents
Facial landmarks for every frame
68 3D points per frame (x, y, z coordinates)
Format: Parquet
Manual annotation using ibug68 template
Format
The landmarks.parquet file contains:
Column
Description
speaker_id
Speaker identifier (1-33)
phrase_id
Phrase identifier (1-10)
frame_id
Frame identifier (e.g., "001", "002")… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-landmarks.welsh-speech-3d-meshes
Welsh Speech Dataset - 3D Facial Meshes
3D facial reconstructions from the Welsh Speech Dataset.
Contents
3D meshes (.obj files) - One per frame
Texture maps (.png files) - Fused left-right stereo images from 3DMD
Captured using 3DMD 6-camera system
~330 zip files (one per speaker-phrase sequence)
File Structure
Files are organized as zip archives in the meshes/ directory, one zip per speaker-phrase sequence:
meshes/
├── speaker_01_phrase_01.zip
├──… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-3d-meshes.welsh-speech-dataset
Welsh Speech Dataset
A multimodal dataset of 33 speakers producing 10 Welsh phrases, captured using 3DMD technology with audio and dense facial landmark annotations.
Dataset Overview
Speakers: 33 participants
Phrases: 10 Welsh phrases per speaker
Sequences: ~330 (33 speakers x 10 phrases)
Modalities:
Audio recordings (.wav)
3D facial reconstructions (.obj meshes + texture maps)
68-point facial landmarks (ibug68 template)
Fluency Scores: Each phrase rated 0-5 (5 =… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-dataset.welsh-speech-dataset
Welsh Speech Dataset
A multimodal dataset of 33 speakers producing 10 Welsh phrases, captured using 3DMD technology with audio and dense facial landmark annotations.
Dataset Overview
Speakers: 33 participants
Phrases: 10 Welsh phrases per speaker
Sequences: ~330 (33 speakers x 10 phrases)
Modalities:
Audio recordings (.wav)
3D facial reconstructions (.obj meshes + texture maps)
68-point facial landmarks (ibug68 template)
Fluency Scores: Each phrase rated 0-5… See the full description on the dataset page: https://huggingface.co/datasets/Pedramebd/welsh-speech-dataset.Welsh_sentiment
