datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Truebones-ZOO-Annotations
Truebones ZOO Annotations
Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for
Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds,
reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly
30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds.
The motion files themselves are not in this repository. Truebones ZOO is a commercial
library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.Truebones-ZOO-Annotations
Truebones ZOO Annotations
Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for
Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds,
reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly
30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds.
The motion files themselves are not in this repository. Truebones ZOO is a commercial
library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Truebones-ZOO-Annotations.JQL-LLM-Edu-Annotations
📚 JQL Educational Quality Annotations from LLMs
This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper.
📝 Dataset Summary
Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs:
Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.seamless-interaction-jefferson-annotations
Seamless Interaction Jefferson-Style Annotations
An automatic, turn-oriented annotation layer for the
Meta Seamless Interaction Dataset.
It compares the dataset's traditional transcript with an ASR-derived
Jefferson-style condition and supplies speech-act, communicative-purpose,
interactional-signal, alignment, and quality fields.
This is a derived noncommercial research dataset. It does not redistribute
the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.vast27m_annotations
VAST-27M Annotations Dataset
This dataset contains annotations from the VAST-27M dataset, originally created for the paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset" by Chen et al. (2024).
Original Source
This dataset is derived from the VAST-27M dataset, which was created by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. The original dataset and more… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/vast27m_annotations.SWE-bench_Verified_With_Annotationscossmos-annotations-dbannotations
The AI Observatory
A public measurement platform aggregating real-world AI conversations from seven sources under a unified 145-feature taxonomy.
This dataset card accompanies paper The AI Observatory: A Public Measure of Real-World AI Use.
📄 Paper: [anonymous OpenReview link]
📊 Dashboard: https://project-ai-observatory.vercel.app/
💾 Anonymous code: https://anonymous.4open.science/r/ai-observatory/README.md
TL;DR
23,158 conversations, 85,633 turns, ~5,000… See the full description on the dataset page: https://huggingface.co/datasets/aiobservatory/annotations.CEFR-Sentence-Level-Annotations
Dataset Card for Dataset Name
17k english sentences annotated by english education professionals. Original repo for CEFR-SP is located at this repo
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/edesaras/CEFR-Sentence-Level-Annotations.TurkWeb-Edu-AnnotationsV3
TurkWeb-Edu V3
Model: Qwen/Qwen3-30B-A3B-Instruct-2507
Format: Structured JSON (vLLM 0.15.0)
emboss-roof-annotations
Emboss 3D Roof Reference Annotations
Manual 3D reference meshes and editable annotations for Swiss and Brazilian buildings, prepared for the evaluation and parameter tuning of Emboss. The annotations describe building and roof geometry, including roof superstructures.
Emboss source code
3dlabel annotation tool
Emboss segmentation model
Example reference annotation in 3dlabel: annotated mesh and LiDAR points (Figure D.1(a) in the paper).
gpqa-diamond-annotations
GPQA Diamond Dataset
This dataset contains filtered JSONL files of human annotations on question specificity, answer uniqueness, answer matching to the ground truth for different models for the GPQA Diamond dataset.
The dataset was annotated by two human graders. It contains 198 (original size) * 2 = 396 rows as each rows is repeated twice (one for each human).
A human grader given the question, actual answer and model response, has to answer whether the response matches the… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/gpqa-diamond-annotations.kwext-bilibili-video-title-annotations
KwExt Bilibili Video Title Annotations
This dataset is a model-assisted annotation set for the KwExt keyword
extraction project. The current snapshot contains 5,000 Chinese Bilibili
video titles from annotation stages video_title_zh_001 through
video_title_zh_005, with 1,000 records in each stage.
The release is intended for early experiments with:
extracting title-grounded keywords and ranking their importance;
broad semantic tags for retrieval and RAG metadata;
dense tag… See the full description on the dataset page: https://huggingface.co/datasets/Himpq/kwext-bilibili-video-title-annotations.bigos-annotationsegolongqa-synth-annotations
EgoLongQA synthetic MCQs, teacher traces and annotation outputs
Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026
EgoLongQA ≤2B track, other than the distillation set (which lives in
infinitylogesh/egolongqa-junior-distill).
⚠️ Read this before counting rows
The synthetic set is 943 questions over 408 videos, and it is stored two ways:
file
rows
shape
training_sets/train_synth_v3.jsonl
943
flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.mania-pattern-annotations
osu!mania pattern annotations
Snapshot v3 uses publication schema beatmap-lens-annotations version 4.
This snapshot contains 592 human judgments, 4403 agent judgments, and 545 source identities (annotation and required calibration sources).
Export implementation: GitHub commit 647009ab60ed.
v3 release scope
The default human layer contains the current effective human observations, with
explicit High/Low confidence where recorded. The opt-in machine layer contains… See the full description on the dataset page: https://huggingface.co/datasets/sed-i/mania-pattern-annotations.4k-video-annotations
4K Video Annotations — Shot Segmentation and Camera Motion
This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties.
The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.example_10kbp_human_annotationsTurkWeb-Edu-AnnotationsV3
TurkWeb-Edu V3
Model: Qwen/Qwen3-30B-A3B-Instruct-2507
Format: Structured JSON (vLLM 0.15.0)
wearable-agent-trajectory-annotations
Wearable Agent Trajectory Annotation Dataset
Dataset Summary
50 wearable agent trajectories annotated by 5 LLM-simulated annotator personas
using the agenteval-schema-v1 JSON schema, across two calibration phases
(500 annotation records total). Designed to benchmark annotation-quality pipelines
for agentic AI systems.
Each trajectory captures a wearable AI agent responding to a real-time sensor event
(health alert, privacy-sensitive context, location trigger… See the full description on the dataset page: https://huggingface.co/datasets/finaspirant/wearable-agent-trajectory-annotations.narrative-gold-annotations
Narrative annotation dataset
Human annotations for three narrative-analysis tasks — setting, agency,
and event relation — over passages sampled from the Dolma corpus.
Annotators & anonymization
Annotator identities are anonymized. Each task has a single gold adjudicator
plus one or more secondary annotators used for double annotation / agreement.
Role
Meaning
gold
The adjudicated / primary label for every released instance.
annotator_1
Second… See the full description on the dataset page: https://huggingface.co/datasets/CLS-Lab/narrative-gold-annotations.example_eval_only_10kb_human_annotationssteam-reviews-constructiveness-binary-label-annotations-1.5k
1.5K Steam Reviews Binary Labeled for Constructiveness
Dataset Summary
This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain.
Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.emolia-voicenet-gemini-annotations
Emolia VoiceNet Gemini Annotations
468,180 dimension-level annotations over 236,613 Emolia speech clips,
each scored 0-6 (0-2 for the content-safety dimension) on one of 57 perceptual
voice / speech dimensions - arousal, valence, brightness, resonance placement, speaking
styles, genuineness, recording quality, and more - by Gemini 3.5 Flash (non-thinking,
temperature 0). This repository ships the annotations, audio provenance, per-dimension
statistics, and the full scoring… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-voicenet-gemini-annotations.narrative-llm-annotations
NarraDolma LLM-Labeled — Distillation Set
The intermediate, LLM-labeled dataset that bridges the small human gold set and the
full NarraDolma corpus. It contains 5,000 passages sampled from
Dolma and labeled by Gemma across
all 11 narrative dimensions, stratified by source and topic to preserve the original
distribution. These labels are the knowledge-distillation training set used to
train NarraBert.
Paper: arXiv:2606.19468
Collection: Narratives in LLM Pretraining Data… See the full description on the dataset page: https://huggingface.co/datasets/CLS-Lab/narrative-llm-annotations.libero-10-execution-annotations
LIBERO_10 Execution Annotations
This repository contains annotation-only metadata for episode executions from the upstream dataset lerobot/libero_10.
It does not redistribute the original dataset, videos, states, actions, or frames.
Each row maps back to the upstream source through:
repo_id / source_repo_id
episode_index / source_episode_index
What Is Included
execution summaries and state descriptions
coarse and detailed episode-level summaries/action labels with… See the full description on the dataset page: https://huggingface.co/datasets/DaivdYuan/libero-10-execution-annotations.wrbench-human-annotations
WRBench Human Annotations
This dataset contains the human comparison labels used to validate WRBench's
automatic evaluation metrics.
Version Update: 2026-07-07
We updated the release after rechecking videos that changed during benchmark
maintenance. The release now includes:
1,741 clean comparison rows.
4,302 individual human judgments.
585 newly rechecked current-benchmark comparisons, each reviewed by three
annotators.
Majority-label summaries for the newly… See the full description on the dataset page: https://huggingface.co/datasets/WRBench/wrbench-human-annotations.corral-reasoning-annotations
Corral – Reasoning Annotations
LLM epistemic annotations over Corral traces where the annotator judged that the agents do not reason scientifically
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains annotated evaluation traces with LLM-generated epistemic annotations across the Corral benchmark.
The dataset is exposed as three model-specific… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-reasoning-annotations.Agentglass-swerebench-annotations
Dataset Summary
SWE-rebench-OpenHands-Trajectories is a dataset of multi-turn agent trajectories for software engineering tasks, collected
using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.54.0) agent scaffolding.
This dataset captures complete agent execution traces as they attempt to resolve real GitHub issues from
nebius/SWE-rebench.
Each trajectory contains the agent's step-by-step reasoning, actions, and environmental observations.
Metric… See the full description on the dataset page: https://huggingface.co/datasets/Abdu07/Agentglass-swerebench-annotations.ConceptARC_Rule_Annotations
ConceptARC-style rule annotations
This release bundles grid-style tasks with natural-language rules written by people or proposed by models, human ratings of how well those rules match the task, and automatic checks of whether each model’s output grid is correct. The is_correct and err fields capture only that grid check, not whether a human endorsed the rule wording. Each row is one attempt on one test case.
On the Hugging Face Hub, the dataset card YAML exposes two subsets (see… See the full description on the dataset page: https://huggingface.co/datasets/AIHumanAbstraction/ConceptARC_Rule_Annotations.
