CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ce-amtic /ProcVQA-20M-annotationsgated ProcVQA-20M Annotations Project Page | arXiv | Code | Model | Media This repository contains the text annotations for the ProcVQA-20M dataset. The full image files are hosted separately on ProcVQA-20M-media. Overview This dataset is constructed from over 26 embodied datasets, comprising: 20M QA pairs for training 330K original trajectories 50M annotated frames from ~5,000 hours of manipulation data 200+ different tasks Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/ce-amtic/ProcVQA-20M-annotations.image10K<n<100K0 likes1.2k downloads4mo agoHugging Face02JQL-AI /JQL-LLM-Edu-Annotations 📚 JQL Educational Quality Annotations from LLMs This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.tabular10M<n<100M2 likes1k downloads1y agoHugging Face03Bofeee5675 /GUI-Net-1M-relative-annotationstext100K<n<1M0 likes691 downloads11mo agoHugging Face04Bofeee5675 /GUI-Net-1M-absolute-annotationstext100K<n<1M2 likes305 downloads11mo agoHugging Face05aiobservatory /annotations The AI Observatory A public measurement platform aggregating real-world AI conversations from seven sources under a unified 145-feature taxonomy. This dataset card accompanies paper The AI Observatory: A Public Measure of Real-World AI Use. 📄 Paper: [anonymous OpenReview link] 📊 Dashboard: https://project-ai-observatory.vercel.app/ 💾 Anonymous code: https://anonymous.4open.science/r/ai-observatory/README.md TL;DR 23,158 conversations, 85,633 turns, ~5,000… See the full description on the dataset page: https://huggingface.co/datasets/aiobservatory/annotations.tabulartext-classification100K<n<1M2 likes298 downloads1mo agoHugging Face06nvidia /SEED-Timeline-Annotations Timeline Annotations for BONES-SEED Humanoid Motion Dataset Dataset Description: This dataset provides additional text description annotations from the BONES-SEED humanoid motion dataset. For each motion, this dataset provides an overview text description of the entire motion at a high level, along with a “timeline” of annotated segments within the motion. Each segment generally contains a single atomic action and is defined by a start time, end time, and text… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SEED-Timeline-Annotations.text100K<n<1M7 likes164 downloads6mo agoHugging Face07nikhilchandak /gpqa-diamond-annotations GPQA Diamond Dataset This dataset contains filtered JSONL files of human annotations on question specificity, answer uniqueness, answer matching to the ground truth for different models for the GPQA Diamond dataset. The dataset was annotated by two human graders. It contains 198 (original size) * 2 = 396 rows as each rows is repeated twice (one for each human). A human grader given the question, actual answer and model response, has to answer whether the response matches the… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/gpqa-diamond-annotations.tabularn<1K1 likes148 downloads1y agoHugging Face08rafmacalaba /data-use-annotations Data-use annotations Public store of keep/drop rulings from the annotation review app (human_labeling/review.html). Files rulings/<annotator>.jsonl — one file per annotator, one JSON object per ruling: key (span UID), ruling (DATA_MENTION keep / NON_MENTION drop), queue (gold / sample), annotator (required, set in the UI), ts. Last write per (queue, key, annotator) wins. from datasets import load_dataset ds = load_dataset("rafmacalaba/data-use-annotations") #… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-annotations.texttext-classificationn<1K0 likes142 downloads12d agoHugging Face09Himpq /kwext-bilibili-video-title-annotations KwExt Bilibili Video Title Annotations This dataset is a model-assisted annotation set for the KwExt keyword extraction project. The current snapshot contains 5,000 Chinese Bilibili video titles from annotation stages video_title_zh_001 through video_title_zh_005, with 1,000 records in each stage. The release is intended for early experiments with: extracting title-grounded keywords and ranking their importance; broad semantic tags for retrieval and RAG metadata; dense tag… See the full description on the dataset page: https://huggingface.co/datasets/Himpq/kwext-bilibili-video-title-annotations.tabulartoken-classification1K<n<10K0 likes127 downloads5d agoHugging Face10ambient-intelligence-labs /egolongqa-synth-annotations EgoLongQA synthetic MCQs, teacher traces and annotation outputs Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026 EgoLongQA ≤2B track, other than the distillation set (which lives in infinitylogesh/egolongqa-junior-distill). ⚠️ Read this before counting rows The synthetic set is 943 questions over 408 videos, and it is stored two ways: file rows shape training_sets/train_synth_v3.jsonl 943 flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.tabularvisual-question-answering1K<n<10K0 likes121 downloads17d agoHugging Face11LianeMarilin /4k-video-annotations 4K Video Annotations — Shot Segmentation and Camera Motion This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties. The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.imagen<1K0 likes110 downloads8d agoHugging Face12BUT-FIT /orca-audio-qa-annotations ORCA Audio QA Annotations Annotation data for training and evaluating ORCA (Open-ended Response Correctness Assessment), a scoring model for audio question-answering tasks. Paper: ORCA: Open-ended Response Correctness Assessment for Audio Question Answering — accepted to TACL 2026 Code & usage: github.com/BUTSpeechFIT/ORCA Pretrained Models: orca-olmo-2-1b-multinomial orca-gemma-3-4b-it-multinomial orca-llama-3.2-3b-it-multinomial Dataset overview ORCA is… See the full description on the dataset page: https://huggingface.co/datasets/BUT-FIT/orca-audio-qa-annotations.texttext-classification100K<n<1M0 likes87 downloads3mo agoHugging Face13JQL-AI /JQL-Human-Edu-Annotations 📚 JQL Multilingual Educational Quality Annotations This dataset provides high-quality human annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Documents: 511 English texts Annotations: 3 human ratings per document (0–5 scale) Translations: Into 35 European languages using DeepL and GPT-4o Purpose: For training and… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-Human-Edu-Annotations.texttext-classification10K<n<100K5 likes70 downloads1y agoHugging Face14thebajajra /muse-trajectory-annotations MUSE trajectory annotations Judge annotations of coding-agent trajectories. Subset: commit-hook Event-sequence annotations of 7,593 transcript windows drawn from 433 complete trajectories of a coding agent working on a git pre-commit-hook task (E1). For each window the judge identifies the earliest concrete workaround opportunity, the earliest rejection of a workaround (labelled normative / instrumental / mixed / unclear), and the earliest later adoption, with… See the full description on the dataset page: https://huggingface.co/datasets/thebajajra/muse-trajectory-annotations.tabular1K<n<10K0 likes68 downloads28d agoHugging Face15MIT-Media-Lab /egotouch-annotations-v1-leftfix egotouch-annotations-v1-leftfix Annotations only. This repository does not contain images or video. This release corrects the left-hand MANO rotation convention in egotouch-annotations-v1. It keeps the original episode structure, instructions, tactile data, and sample index. Item Count Episodes 111,159 Frames and training samples 3,687,389 Repaired left-hand episodes 54,373 Repaired left-hand frames 1,781,845 Unchanged right-only episodes 56,786… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/egotouch-annotations-v1-leftfix.textroboticsn<1K0 likes50 downloads14d agoHugging Face16jamesding0302 /memgen-annotations MemGen Annotations This is the annotation dataset for the paper How Well Does Generative Recommendation Generalize?. The annotations categorize evaluation instances under the leave-one-out protocol: test split uses the last item in the user history sequence as target, val split uses the second-to-last item as target. Columns sample_id: row index within the split in the original dataset. user_id: raw user identifier (join key). master: one of memorization… See the full description on the dataset page: https://huggingface.co/datasets/jamesding0302/memgen-annotations.textother100K<n<1M1 likes48 downloads6mo agoHugging Face17toroe /Dolci-Think-SFT-7B-Propella-Annotationstabular1M<n<10M0 likes39 downloads7mo agoHugging Face18Hannibal52Barca /icl-sarm-annotations ICL SARM Subtask Annotations Per-episode subtask decomposition (name + start/end frame) generated with a VLM-based annotation pipeline (Qwen3-VL-8B-Instruct, ecot-style plan generation + bidirectional grounding), for the two adityx23 ICL robot manipulation datasets: File Source dataset Episodes Tasks icl-dataset_subtasks.jsonl adityx23/icl-dataset (reference set) 3149 36 icl-demo-dataset_subtasks.jsonl adityx23/icl-demo-dataset (query/test set) 285 27… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/icl-sarm-annotations.textrobotics1K<n<10K0 likes36 downloads8d agoHugging Face19Sierkinhane /show-o2-data-annotationstext10K<n<100K1 likes33 downloads1y agoHugging Face20himalaya-ai /nepali-stt-annotationstabularn<1K0 likes33 downloads1mo agoHugging Face21dill-lab /oath-frames-expert-annotationstext1K<n<10K0 likes31 downloads2y agoHugging Face22netprtony /pokemon-cards-image-and-annotationsimage1K<n<10K0 likes31 downloads11mo agoHugging Face23VLAI-AIVN /DAM-QA-annotations DAM-QA Unified Annotations 22,675 question-answer pairs from 6 major VQA benchmarks, unified for the DAM-QA framework. This collection consolidates annotations from InfographicVQA, TextVQA, VQAv2, DocVQA, ChartQA, and ChartQA-Pro into standardized JSONL formats. 📖 Paper: Describe Anything Model for Visual Question Answering on Text-rich Images⚠️ Note: Images not included - obtain from original sources with proper licensing Repository Structure DAM-QA-annotations/… See the full description on the dataset page: https://huggingface.co/datasets/VLAI-AIVN/DAM-QA-annotations.textquestion-answering10K<n<100K0 likes27 downloads1y agoHugging Face24Experimental-Orange /HumanAgencyBench_Human_Annotations Human annotations and LLM judge comparative Dataset Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains 60,000 evaluated AI assistant responses across 6 dimensions of behaviour relevant to human agency support, with both model-based and human annotations. Each example includes evaluations from 4 different frontier LLM models. We also provide… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Human_Annotations.texttext-generation10K<n<100K0 likes26 downloads1y agoHugging Face25cohort-rlwm /Liquid_V1_7B-pico-aurora-vidgen-multiturn-annotationstext100K<n<1M0 likes26 downloads9mo agoHugging Face26abdo1819 /arabic-english-code-switching-review-annotations Review Annotations for Arabic-English Code-Switching Speech This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts. The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index. Coverage and outcomes The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.tabularautomatic-speech-recognition10K<n<100K0 likes25 downloads1mo agoHugging Face27shivani-kerai /vqa_training_annotationstabular100K<n<1M0 likes22 downloads2y agoHugging Face28bluolightning /manga109s-line-annotations Manga109-s Text Line Annotations High-precision, line-level bounding box and polygon annotations for the Manga109-s Dataset, supporting both full manga pages and speech bubble crops. Furigana is not labeled and is almost entirely excluded from line labels. Includes 8-point oriented polygons for slanted/rotated text lines. The annotation process is documented in METHODOLOGY.md (WIP). Notice: This dataset contains zero dialogue text and zero images. It requires your own local… See the full description on the dataset page: https://huggingface.co/datasets/bluolightning/manga109s-line-annotations.textimage-text-to-textn<1K0 likes22 downloads3d agoHugging Face29jakequist /kanjiland-silver-annotations Kanjiland — Silver Annotations 9,388 Japanese sentences annotated in the full Kanjiland reading-comprehension format: morpheme segmentation, furigana (ruby on kanji runs), per-token contextual glosses, word groupings, an English sentence translation, and grammar-pattern labels from a closed 120-rule inventory. Part of Kanjiland, a from-scratch Japanese reading-comprehension engine. Format Each line of silver_annotations.jsonl is: {"ja": "<Japanese sentence>"… See the full description on the dataset page: https://huggingface.co/datasets/jakequist/kanjiland-silver-annotations.texttranslation1K<n<10K0 likes18 downloads2mo agoHugging Face30Kevius /sanpo_annotationstextn<1K2 likes17 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.