datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hpltv2-llama33-edu-annotation
HPLT version 2.0 educational annotations
This dataset contains annotations derived from HPLT v2 cleaned samples.
There are 500,000 annotations for each language if the source contains at least 500,000 samples.
We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier.
Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.ProcVQA-20M-annotations
ProcVQA-20M Annotations
Project Page |
arXiv |
Code |
Model |
Media
This repository contains the text annotations for the ProcVQA-20M dataset. The full image files are hosted separately on ProcVQA-20M-media.
Overview
This dataset is constructed from over 26 embodied datasets, comprising:
20M QA pairs for training
330K original trajectories
50M annotated frames from ~5,000 hours of manipulation data
200+ different tasks
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/ce-amtic/ProcVQA-20M-annotations.JQL-LLM-Edu-Annotations
📚 JQL Educational Quality Annotations from LLMs
This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper.
📝 Dataset Summary
Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs:
Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.GUI-Net-1M-relative-annotationsLocateAnything-Data-ShareGPT-AnnotationGUI-Net-1M-absolute-annotationsannotations
The AI Observatory
A public measurement platform aggregating real-world AI conversations from seven sources under a unified 145-feature taxonomy.
This dataset card accompanies paper The AI Observatory: A Public Measure of Real-World AI Use.
📄 Paper: [anonymous OpenReview link]
📊 Dashboard: https://project-ai-observatory.vercel.app/
💾 Anonymous code: https://anonymous.4open.science/r/ai-observatory/README.md
TL;DR
23,158 conversations, 85,633 turns, ~5,000… See the full description on the dataset page: https://huggingface.co/datasets/aiobservatory/annotations.SenseNova-Vision-Corpus-50M-annotationSEED-Timeline-Annotations
Timeline Annotations for BONES-SEED Humanoid Motion Dataset
Dataset Description:
This dataset provides additional text description annotations from the BONES-SEED humanoid motion dataset. For each motion, this dataset provides an overview text description of the entire motion at a high level, along with a “timeline” of annotated segments within the motion. Each segment generally contains a single atomic action and is defined by a start time, end time, and text… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SEED-Timeline-Annotations.gpqa-diamond-annotations
GPQA Diamond Dataset
This dataset contains filtered JSONL files of human annotations on question specificity, answer uniqueness, answer matching to the ground truth for different models for the GPQA Diamond dataset.
The dataset was annotated by two human graders. It contains 198 (original size) * 2 = 396 rows as each rows is repeated twice (one for each human).
A human grader given the question, actual answer and model response, has to answer whether the response matches the… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/gpqa-diamond-annotations.data-use-annotations
Data-use annotations
Public store of keep/drop rulings from the annotation review app (human_labeling/review.html).
Files
rulings/<annotator>.jsonl — one file per annotator, one JSON object per ruling: key (span UID), ruling (DATA_MENTION keep / NON_MENTION drop), queue (gold / sample), annotator (required, set in the UI), ts. Last write per (queue, key, annotator) wins.
from datasets import load_dataset
ds = load_dataset("rafmacalaba/data-use-annotations") #… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-annotations.kwext-bilibili-video-title-annotations
KwExt Bilibili Video Title Annotations
This dataset is a model-assisted annotation set for the KwExt keyword
extraction project. The current snapshot contains 5,000 Chinese Bilibili
video titles from annotation stages video_title_zh_001 through
video_title_zh_005, with 1,000 records in each stage.
The release is intended for early experiments with:
extracting title-grounded keywords and ranking their importance;
broad semantic tags for retrieval and RAG metadata;
dense tag… See the full description on the dataset page: https://huggingface.co/datasets/Himpq/kwext-bilibili-video-title-annotations.egolongqa-synth-annotations
EgoLongQA synthetic MCQs, teacher traces and annotation outputs
Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026
EgoLongQA ≤2B track, other than the distillation set (which lives in
infinitylogesh/egolongqa-junior-distill).
⚠️ Read this before counting rows
The synthetic set is 943 questions over 408 videos, and it is stored two ways:
file
rows
shape
training_sets/train_synth_v3.jsonl
943
flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.CCI3-HQ-Annotation-Benchmark
CCI3-HQ-Annotation-Benchmark
These 14k samples were randomly extracted from a large corpus of Chinese texts, containing both the original text and corresponding labels. They can be used to evaluate the quality of Chinese corpora.
Citation
If you use this benchmark or the CCI3-HQ dataset, please cite:
@misc{wang2024cci30hqlargescalechinesedataset,
title={CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CCI3-HQ-Annotation-Benchmark.4k-video-annotations
4K Video Annotations — Shot Segmentation and Camera Motion
This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties.
The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.video_annotation_pipeline
👁️ Semi-Automatic Video Annotation Pipeline
📝 Description
Video-ChatGPT introduces the VideoInstruct100K dataset, which employs a semi-automatic annotation pipeline to generate 75K instruction-tuning QA pairs. To address the limitations of this annotation process, we present VCG+112K dataset developed through an improved annotation pipeline. Our approach improves the accuracy and quality of instruction tuning pairs by improving keyframe extraction, leveraging SoTA… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/video_annotation_pipeline.orca-audio-qa-annotations
ORCA Audio QA Annotations
Annotation data for training and evaluating ORCA (Open-ended Response Correctness Assessment), a scoring model for audio question-answering tasks.
Paper: ORCA: Open-ended Response Correctness Assessment for Audio Question Answering — accepted to TACL 2026
Code & usage: github.com/BUTSpeechFIT/ORCA
Pretrained Models:
orca-olmo-2-1b-multinomial
orca-gemma-3-4b-it-multinomial
orca-llama-3.2-3b-it-multinomial
Dataset overview
ORCA is… See the full description on the dataset page: https://huggingface.co/datasets/BUT-FIT/orca-audio-qa-annotations.normistral-fluency-annotationManual fluency annotations for Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
Citation
@misc{samuel2025fluentalignmentdisfluentjudges,
title={Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages},
author={David Samuel and Lilja Øvrelid and Erik Velldal and Andrey Kutuzov},
year={2025},
eprint={2512.08777},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/ltg/normistral-fluency-annotation.JQL-Human-Edu-Annotations
📚 JQL Multilingual Educational Quality Annotations
This dataset provides high-quality human annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper.
📝 Dataset Summary
Documents: 511 English texts
Annotations: 3 human ratings per document (0–5 scale)
Translations: Into 35 European languages using DeepL and GPT-4o
Purpose: For training and… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-Human-Edu-Annotations.muse-trajectory-annotations
MUSE trajectory annotations
Judge annotations of coding-agent trajectories.
Subset: commit-hook
Event-sequence annotations of 7,593 transcript windows drawn from 433 complete
trajectories of a coding agent working on a git pre-commit-hook task (E1). For
each window the judge identifies the earliest concrete workaround opportunity,
the earliest rejection of a workaround (labelled normative / instrumental /
mixed / unclear), and the earliest later adoption, with… See the full description on the dataset page: https://huggingface.co/datasets/thebajajra/muse-trajectory-annotations.egotouch-annotations-v1-leftfix
egotouch-annotations-v1-leftfix
Annotations only. This repository does not contain images or video.
This release corrects the left-hand MANO rotation convention in
egotouch-annotations-v1.
It keeps the original episode structure, instructions, tactile data, and sample index.
Item
Count
Episodes
111,159
Frames and training samples
3,687,389
Repaired left-hand episodes
54,373
Repaired left-hand frames
1,781,845
Unchanged right-only episodes
56,786… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/egotouch-annotations-v1-leftfix.memgen-annotations
MemGen Annotations
This is the annotation dataset for the paper How Well Does Generative Recommendation Generalize?.
The annotations categorize evaluation instances under the leave-one-out protocol:
test split uses the last item in the user history sequence as target,
val split uses the second-to-last item as target.
Columns
sample_id: row index within the split in the original dataset.
user_id: raw user identifier (join key).
master: one of memorization… See the full description on the dataset page: https://huggingface.co/datasets/jamesding0302/memgen-annotations.Dolci-Think-SFT-7B-Propella-Annotationssotopia-rl-reward-annotation
Sotopia-RL: Reward Design for Social Intelligence Dataset
This repository contains the dataset and related resources for the paper Sotopia-RL: Reward Design for Social Intelligence.
Sotopia-RL proposes a novel framework that refines coarse episode-level feedback into utterance-level, multi-dimensional rewards. This enables more effective training of socially intelligent agents through reinforcement learning, particularly addressing challenges like partial observability and… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/sotopia-rl-reward-annotation.Sinhala_Annotation_Dataset
Sinhala Named Entity Recognition (NER) Dataset - 85,000 Annotations
Dataset Description
This is a high-quality Named Entity Recognition (NER) dataset for the Sinhala language, consisting of approximately 85,000 annotations. The dataset was manually curated and annotated by a team of three students to support NLP research for low-resource languages.
The data is sourced from diverse domains, including social media comments, news articles, and public domain texts, capturing… See the full description on the dataset page: https://huggingface.co/datasets/kasunUdayanga/Sinhala_Annotation_Dataset.icl-sarm-annotations
ICL SARM Subtask Annotations
Per-episode subtask decomposition (name + start/end frame) generated with a VLM-based
annotation pipeline (Qwen3-VL-8B-Instruct, ecot-style plan generation +
bidirectional grounding), for the two adityx23 ICL robot manipulation datasets:
File
Source dataset
Episodes
Tasks
icl-dataset_subtasks.jsonl
adityx23/icl-dataset (reference set)
3149
36
icl-demo-dataset_subtasks.jsonl
adityx23/icl-demo-dataset (query/test set)
285
27… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/icl-sarm-annotations.wikipedia-22-12-en-annotationshow-o2-data-annotationspokemon-cards-image-and-annotations
