datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lasa1m-annotate-part-12lasa1m-annotate-part-14lasa1m-annotate-part-16lasa1m-annotate-part-15lasa1m-annotate-part-06lasa1m-annotate-part-07lasa1m-annotate-part-05lasa1m-annotate-part-13lasa1m-annotate-part-09lasa1m-annotate-part-10lasa1m-annotate-part-11lasa1m-annotate-part-08fitcheck-annotate-datasetlasa1m-annotate-part-01pico-robotics-annotated
Pico Robotics Dataset · Annotated Edition
Egocentric multimodal capture from a Pico VR headset + custom tracker rig — Annotated tier
Builds on the Advanced edition by adding coarse action segmentation. Every sequence is divided
into labelled temporal segments, so the data can be used directly for action recognition,
temporal segmentation, and behaviour-understanding tasks without an annotation pass of your own.
🔒 This is a gated dataset. Access requests are reviewed manually;… See the full description on the dataset page: https://huggingface.co/datasets/skycn110/pico-robotics-annotated.lasa1m-annotate-part-04lasa1m-annotate-part-03lasa1m-annotate-part-02muaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']
وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.mualem-recitations-annotatedwildchat_creative_writing_annotated_10kMalaysian-Emilia-annotated
Malaysian Emilia Annotated
Annotate Malaysian-Emilia using Data-Speech pipeline.
Malaysian Youtube
Originally from malaysia-ai/crawl-youtube
Total 3168.8 hours.
Gender prediction, filtered-24k_processed_24k_gender.zip
Language prediction, filtered-24k_processed_language.zip
Force alignment.
Post cleaned to 24k and 44k sampling rates,
24k, filtered-24k_processed_24k.zip
44k, filtered-24k_processed_44k.zip
Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.robocasa_pretrain_human300_v4_annotated5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.robot0_agentview_left": {
"dtype": "video",
"shape": [
256,
256,
3
],
"names": [
"height",
"width",
"channel"
],
"video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/robocasa_pretrain_human300_v4_annotated5.GAIA-annotateddronescapes2_annotated_train_set
Dataset Card for DroneScapes2 (annotated train set)
This is a FiftyOne dataset with 218 samples. It's a subset of this split from the original repo.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/dronescapes2_annotated_train_set.laion-voice-profiles-annotated
LAION Voice Profiles — Annotated
Authors: Christoph Schuhmann and LAION.
28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from
500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting
conditions. Every
utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads,
vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a
768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.fineweb_annotatedannotated-3DGS-artifacts
Puzzle Similarity
Project page | Paper | Code
This repository contains the dataset presented in the ICCV 2025 paper "Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene Reconstructions"Authors: Nicolai Hermann, Jorge Condor, and Piotr Didyk
Dataset Description
The Dataset consists of 36 hand-selected 3D Gaussian Splatting renderings containing common reconstruction artefacts, (aligned) ground truths, human-annotated… See the full description on the dataset page: https://huggingface.co/datasets/nihermann/annotated-3DGS-artifacts.RPCDr1_annotated_aime
