datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
co3d-annotations
CO3D Annotations (Derived)
This dataset is derived from the CO3D Dataset.
For Test Set ONLY,
Preprocessed for VGGT input
License
Original License: Creative Commons Attribution-NonCommercial 4.0 (CC BY-NC 4.0)
Non-commercial use only.
Attribution required: This dataset is derived from the CO3D dataset © Meta, licensed under CC BY-NC 4.0.
propella-annotations
This dataset contains document annotations produced with propella-1-4b, a small multilingual LLM that annotates text documents across six categories: core content, classification, quality & value, audience & purpose, safety & compliance, and geographic relevance. The annotations can be used to filter, select, and curate LLM training data at scale.
Properties
Each document is annotated across 18 properties organized into six categories:
Category
Property
Description… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/propella-annotations.laions_got_talent_enhanced_flash_annotations_and_long_captionsTruebones-ZOO-Annotations
Truebones ZOO Annotations
Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for
Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds,
reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly
30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds.
The motion files themselves are not in this repository. Truebones ZOO is a commercial
library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.laions_got_talent_enhanced_just_flash_annotationsbeat2-additional-annotations
BEAT2 Official Release + Additional Annotations
This is a fork of H-Liu1997/BEAT2
that adds annotations contributed by the
RAG-Gesture (CVPR 2025)
and MIBURI (CVPR 2026) projects.
The base BEAT2-English data (motion, audio, TextGrids, semantic labels,
pretrained motion-autoencoder weights) is inherited verbatim from upstream;
the additional annotations from RAG-Gesture and MIBURI are pushed on top.
Citations
If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.VideoChat3-Training-Data-Annotations
VideoChat3-Stage3-Training-Data
This repository includes all annotation files used across the four training stages of VideoChat3, from Stage 0 to Stage 3.
You can refer to the provided source-data links to download videos, images, and other multimedia data for training. In videochat3_data_annotations, we also provide a source field to indicate the source dataset for each entry.
To facilitate Stage 3 training reproduction using the high-quality open-source datasets we collected… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-Training-Data-Annotations.Emilia-with-Emotion-Annotations
Dataset Card for Emilia with Emotion Annotations
Dataset Description
This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?".
The annotations include 54 scores for each sample, covering a wide range… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emilia-with-Emotion-Annotations.gsd-humaneval-annotationsTruebones-ZOO-Annotations
Truebones ZOO Annotations
Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for
Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds,
reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly
30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds.
The motion files themselves are not in this repository. Truebones ZOO is a commercial
library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Truebones-ZOO-Annotations.dcvlm_pool_small_annotations
DCVLM-Pool (small) — per-sample annotations
Every filtering annotation we computed for the small data pool of our
DataComp-VLM paper: image quality, image–text alignment, language ID,
text-quality classifiers, multimodal perplexity, decontamination scores and more — up to 167 fields per
sample (180 distinct fields overall), for all 120,940,134 samples across 166 source datasets.
These are the raw annotations, not a filtered dataset. They are the inputs our curation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_small_annotations.ProcVQA-20M-annotations
ProcVQA-20M Annotations
Project Page |
arXiv |
Code |
Model |
Media
This repository contains the text annotations for the ProcVQA-20M dataset. The full image files are hosted separately on ProcVQA-20M-media.
Overview
This dataset is constructed from over 26 embodied datasets, comprising:
20M QA pairs for training
330K original trajectories
50M annotated frames from ~5,000 hours of manipulation data
200+ different tasks
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/ce-amtic/ProcVQA-20M-annotations.open-greek-corpus-annotations
Open Greek Corpus Annotations
Token-level linguistic annotations for the
Open Greek Corpus:
lemma, part of speech (UD UPOS), and morphology (UD features) for every
served token. Three provenance classes, never confused thanks to per-token
provenance and confidence tiers: gold treebank annotations where an openly
licensed MANUAL treebank covers a work (GLAUx's treebank layers, MACULA
Greek for the NT), GLAUx's own automatic annotation as the middle auto:
class, and model… See the full description on the dataset page: https://huggingface.co/datasets/ciscoriordan/open-greek-corpus-annotations.JQL-LLM-Edu-Annotations
📚 JQL Educational Quality Annotations from LLMs
This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper.
📝 Dataset Summary
Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs:
Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.Emilia-with-Emotion-Annotations4seamless-interaction-jefferson-annotations
Seamless Interaction Jefferson-Style Annotations
An automatic, turn-oriented annotation layer for the
Meta Seamless Interaction Dataset.
It compares the dataset's traditional transcript with an ASR-derived
Jefferson-style condition and supplies speech-act, communicative-purpose,
interactional-signal, alignment, and quality fields.
This is a derived noncommercial research dataset. It does not redistribute
the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.python-edu-annotations
Annotations for 📚 Python-Edu classifier
This dataset contains the annotations used for training Python-Edu educational quality classifier. We prompt Llama-3-70B-Instruct to score python programs from StarCoderData based on their educational value.
Note: the dataset contains the Python program, the prompt (using the first 1000 characters of the program) and the scores but it doesn't contain the full Llama 3 generation.
Emilia-with-Emotion-Annotations5GUI-Net-1M-relative-annotationsvast27m_annotations
VAST-27M Annotations Dataset
This dataset contains annotations from the VAST-27M dataset, originally created for the paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset" by Chen et al. (2024).
Original Source
This dataset is derived from the VAST-27M dataset, which was created by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. The original dataset and more… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/vast27m_annotations.probe-tip-annotations-datapdfQA-Annotations
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research.
This repository contains the pdfQA-Annotations dataset, which provides only the QA annotations and metadata for the pdfQA-Benchmark.
It is intended for lightweight experimentation, modeling, and evaluation without requiring access to large document files.
Relationship to the Full pdfQA… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Annotations.od-syn-page-annotations-com
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script.
Note: this version image are compressed.
Raw version 📁 Repository: Hugging Face Datasets
📋 Dataset Summary
Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.fineweb-edu-llama3-annotations
Annotations for 📚 FineWeb-Edu classifier
This dataset contains the annotations used for training 📚 FineWeb-Edu educational quality classifier. We prompt Llama-3-70B-Instruct to score web pages from 🍷 FineWeb based on their educational value.
Note: the dataset contains the FineWeb text sample, the prompt (using the first 1000 characters of the text sample) and the scores but it doesn't contain the full Llama 3 generation.
SWE-bench_Verified_With_Annotationsod-syn-page-annotations
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis , visual document understanding , OCR fine-tuning, and related tasks specifically for Dhivehi script.
📋 Dataset Summary
Total Examples: ~58,738
Image Content: Synthetic Dhivehi documents generated to simulate real-world layouts… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations.Emilia-with-Emotion-Annotations3merge_annotations_self_and_swegymcossmos-annotations-dbEmilia-with-Emotion-Annotations2
