CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LumiOpen /hpltv2-llama33-edu-annotation HPLT version 2.0 educational annotations This dataset contains annotations derived from HPLT v2 cleaned samples. There are 500,000 annotations for each language if the source contains at least 500,000 samples. We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier. Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.tabular10M<n<100M3 likes1.4k downloads1y agoHugging Face02JQL-AI /JQL-LLM-Edu-Annotations 📚 JQL Educational Quality Annotations from LLMs This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.tabular10M<n<100M2 likes1k downloads1y agoHugging Face03aiobservatory /annotations The AI Observatory A public measurement platform aggregating real-world AI conversations from seven sources under a unified 145-feature taxonomy. This dataset card accompanies paper The AI Observatory: A Public Measure of Real-World AI Use. 📄 Paper: [anonymous OpenReview link] 📊 Dashboard: https://project-ai-observatory.vercel.app/ 💾 Anonymous code: https://anonymous.4open.science/r/ai-observatory/README.md TL;DR 23,158 conversations, 85,633 turns, ~5,000… See the full description on the dataset page: https://huggingface.co/datasets/aiobservatory/annotations.tabulartext-classification100K<n<1M2 likes294 downloads1mo agoHugging Face04nikhilchandak /gpqa-diamond-annotations GPQA Diamond Dataset This dataset contains filtered JSONL files of human annotations on question specificity, answer uniqueness, answer matching to the ground truth for different models for the GPQA Diamond dataset. The dataset was annotated by two human graders. It contains 198 (original size) * 2 = 396 rows as each rows is repeated twice (one for each human). A human grader given the question, actual answer and model response, has to answer whether the response matches the… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/gpqa-diamond-annotations.tabularn<1K1 likes148 downloads1y agoHugging Face05Himpq /kwext-bilibili-video-title-annotations KwExt Bilibili Video Title Annotations This dataset is a model-assisted annotation set for the KwExt keyword extraction project. The current snapshot contains 5,000 Chinese Bilibili video titles from annotation stages video_title_zh_001 through video_title_zh_005, with 1,000 records in each stage. The release is intended for early experiments with: extracting title-grounded keywords and ranking their importance; broad semantic tags for retrieval and RAG metadata; dense tag… See the full description on the dataset page: https://huggingface.co/datasets/Himpq/kwext-bilibili-video-title-annotations.tabulartoken-classification1K<n<10K0 likes140 downloads5d agoHugging Face06ambient-intelligence-labs /egolongqa-synth-annotations EgoLongQA synthetic MCQs, teacher traces and annotation outputs Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026 EgoLongQA ≤2B track, other than the distillation set (which lives in infinitylogesh/egolongqa-junior-distill). ⚠️ Read this before counting rows The synthetic set is 943 questions over 408 videos, and it is stored two ways: file rows shape training_sets/train_synth_v3.jsonl 943 flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.tabularvisual-question-answering1K<n<10K0 likes127 downloads17d agoHugging Face07LianeMarilin /4k-video-annotations 4K Video Annotations — Shot Segmentation and Camera Motion This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties. The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.imagen<1K0 likes120 downloads8d agoHugging Face08superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes112 downloads5mo agoHugging Face09thebajajra /muse-trajectory-annotations MUSE trajectory annotations Judge annotations of coding-agent trajectories. Subset: commit-hook Event-sequence annotations of 7,593 transcript windows drawn from 433 complete trajectories of a coding agent working on a git pre-commit-hook task (E1). For each window the judge identifies the earliest concrete workaround opportunity, the earliest rejection of a workaround (labelled normative / instrumental / mixed / unclear), and the earliest later adoption, with… See the full description on the dataset page: https://huggingface.co/datasets/thebajajra/muse-trajectory-annotations.tabular1K<n<10K0 likes56 downloads28d agoHugging Face10toroe /Dolci-Think-SFT-7B-Propella-Annotationstabular1M<n<10M0 likes39 downloads7mo agoHugging Face11MongoDB /wikipedia-22-12-en-annotationtabular10K<n<100K0 likes34 downloads2y agoHugging Face12abdo1819 /arabic-english-code-switching-review-annotations Review Annotations for Arabic-English Code-Switching Speech This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts. The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index. Coverage and outcomes The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.tabularautomatic-speech-recognition10K<n<100K0 likes29 downloads1mo agoHugging Face13Flaglab /esnlir-annotation-candidates ESNLIR — confidence-stratified annotation candidates The 2,664 candidate pairs sent to human annotators, each carrying the model confidence score and the stratum it was drawn from. This is the pool before annotation; the labelled result is Flaglab/esnlir-al-annotated-test. Active Learning for Spanish Natural Language Inference on a Heterogeneous Multi-Domain Corpus Diego Ortiz, Johan R. Portela, Ruben Manrique — Universidad de los Andes, Bogotá Advances in Artificial… See the full description on the dataset page: https://huggingface.co/datasets/Flaglab/esnlir-annotation-candidates.tabulartext-classification1K<n<10K0 likes23 downloads2mo agoHugging Face14shivani-kerai /vqa_training_annotationstabular100K<n<1M0 likes22 downloads2y agoHugging Face15himalaya-ai /nepali-stt-annotationstabularn<1K0 likes22 downloads1mo agoHugging Face16kevineen /Tanuki-Phase2-annotation-datasettabulartext-classificationn<1K0 likes17 downloads2y agoHugging Face17monodox /carnatic-gamaka-annotationstabularn<1K0 likes13 downloads5mo agoHugging Face18te-pa /systems-observatory-annotations Te Pā Systems Observatory · Community Annotations Public archive of community-submitted rankings of interventions against Donella Meadows' twelve leverage points, in the Aotearoa New Zealand political-economy context. Ranked by kaimahi (community members) through the Systems Observatory dashboard. Built under the Māori Data Governance Model of Te Kāhui Raraunga. Public, aggregated, non-personal data only. Schema (JSON Lines) Each row in annotations.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/te-pa/systems-observatory-annotations.tabularn<1K0 likes12 downloads2mo agoHugging Face19rsepulvedat /Annotation_testtabularn<1K0 likes9 downloads1y agoHugging Face20laion /unsupervised_peoples_speech_with_annotations1tabularn<1K0 likes4 downloads1y agoHugging Face21Robinnine05 /Restaurant_Review_LLMs_Annotationtabular10K<n<100K0 likes4 downloads4mo agoHugging Face22gohsyi /saferlhf-iter1-annotation-gemma-2-2b-sft.jsonltabular10K<n<100K0 likes3 downloads2y agoHugging Face23redaMM /educational-ai-agent-small-annotation-depthtabularn<1K0 likes3 downloads1y agoHugging Face24vanquorrr /senvlm-annotation-batchtabularn<1K0 likes3 downloads6mo agoHugging Face25acengnew /obstacle-map-annotations-json Obstacle Map Annotations Dataset Dataset annotating obstacles in a 2D environment for robot navigation and mapping. tabularn<1K0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.