datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JQL-LLM-Edu-Annotations
📚 JQL Educational Quality Annotations from LLMs
This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper.
📝 Dataset Summary
Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs:
Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.annotations
The AI Observatory
A public measurement platform aggregating real-world AI conversations from seven sources under a unified 145-feature taxonomy.
This dataset card accompanies paper The AI Observatory: A Public Measure of Real-World AI Use.
📄 Paper: [anonymous OpenReview link]
📊 Dashboard: https://project-ai-observatory.vercel.app/
💾 Anonymous code: https://anonymous.4open.science/r/ai-observatory/README.md
TL;DR
23,158 conversations, 85,633 turns, ~5,000… See the full description on the dataset page: https://huggingface.co/datasets/aiobservatory/annotations.gpqa-diamond-annotations
GPQA Diamond Dataset
This dataset contains filtered JSONL files of human annotations on question specificity, answer uniqueness, answer matching to the ground truth for different models for the GPQA Diamond dataset.
The dataset was annotated by two human graders. It contains 198 (original size) * 2 = 396 rows as each rows is repeated twice (one for each human).
A human grader given the question, actual answer and model response, has to answer whether the response matches the… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/gpqa-diamond-annotations.kwext-bilibili-video-title-annotations
KwExt Bilibili Video Title Annotations
This dataset is a model-assisted annotation set for the KwExt keyword
extraction project. The current snapshot contains 5,000 Chinese Bilibili
video titles from annotation stages video_title_zh_001 through
video_title_zh_005, with 1,000 records in each stage.
The release is intended for early experiments with:
extracting title-grounded keywords and ranking their importance;
broad semantic tags for retrieval and RAG metadata;
dense tag… See the full description on the dataset page: https://huggingface.co/datasets/Himpq/kwext-bilibili-video-title-annotations.egolongqa-synth-annotations
EgoLongQA synthetic MCQs, teacher traces and annotation outputs
Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026
EgoLongQA ≤2B track, other than the distillation set (which lives in
infinitylogesh/egolongqa-junior-distill).
⚠️ Read this before counting rows
The synthetic set is 943 questions over 408 videos, and it is stored two ways:
file
rows
shape
training_sets/train_synth_v3.jsonl
943
flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.4k-video-annotations
4K Video Annotations — Shot Segmentation and Camera Motion
This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties.
The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.muse-trajectory-annotations
MUSE trajectory annotations
Judge annotations of coding-agent trajectories.
Subset: commit-hook
Event-sequence annotations of 7,593 transcript windows drawn from 433 complete
trajectories of a coding agent working on a git pre-commit-hook task (E1). For
each window the judge identifies the earliest concrete workaround opportunity,
the earliest rejection of a workaround (labelled normative / instrumental /
mixed / unclear), and the earliest later adoption, with… See the full description on the dataset page: https://huggingface.co/datasets/thebajajra/muse-trajectory-annotations.Dolci-Think-SFT-7B-Propella-Annotationsarabic-english-code-switching-review-annotations
Review Annotations for Arabic-English Code-Switching Speech
This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts.
The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index.
Coverage and outcomes
The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.vqa_training_annotationsnepali-stt-annotationscarnatic-gamaka-annotationssystems-observatory-annotations
Te Pā Systems Observatory · Community Annotations
Public archive of community-submitted rankings of interventions against Donella
Meadows' twelve leverage points, in the Aotearoa New Zealand political-economy
context. Ranked by kaimahi (community members) through the
Systems Observatory dashboard.
Built under the Māori Data Governance Model of
Te Kāhui Raraunga. Public, aggregated, non-personal data only.
Schema (JSON Lines)
Each row in annotations.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/te-pa/systems-observatory-annotations.unsupervised_peoples_speech_with_annotations1obstacle-map-annotations-json
Obstacle Map Annotations Dataset
Dataset annotating obstacles in a 2D environment
for robot navigation and mapping.
