CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01commoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes12k downloads10mo agoHugging Face02tanish434 /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.tabular1K<n<10K0 likes3.1k downloads9d agoHugging Face03Linzhan /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Truebones-ZOO-Annotations.tabular1K<n<10K1 likes1.7k downloads15d agoHugging Face04LumiOpen /hpltv2-llama33-edu-annotation HPLT version 2.0 educational annotations This dataset contains annotations derived from HPLT v2 cleaned samples. There are 500,000 annotations for each language if the source contains at least 500,000 samples. We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier. Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.tabular10M<n<100M3 likes1.5k downloads1y agoHugging Face05JQL-AI /JQL-LLM-Edu-Annotations 📚 JQL Educational Quality Annotations from LLMs This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.tabular10M<n<100M2 likes1k downloads1y agoHugging Face06kennethli319 /seamless-interaction-jefferson-annotations Seamless Interaction Jefferson-Style Annotations An automatic, turn-oriented annotation layer for the Meta Seamless Interaction Dataset. It compares the dataset's traditional transcript with an ASR-derived Jefferson-style condition and supplies speech-act, communicative-purpose, interactional-signal, alignment, and quality fields. This is a derived noncommercial research dataset. It does not redistribute the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.tabularautomatic-speech-recognition100K<n<1M0 likes873 downloads2mo agoHugging Face07it-just-works /vast27m_annotations VAST-27M Annotations Dataset This dataset contains annotations from the VAST-27M dataset, originally created for the paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset" by Chen et al. (2024). Original Source This dataset is derived from the VAST-27M dataset, which was created by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. The original dataset and more… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/vast27m_annotations.tabular10M<n<100M1 likes689 downloads2y agoHugging Face08commoncrawl /gneissweb-annotation-host-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-host-testing-v1.tabular100M<n<1B1 likes554 downloads10mo agoHugging Face09VR-VLA /VR-egodex-annotation-converted-v6.0 VR-egodex-annotation-converted-v6.0 EgoDex converted from LeRobot v2.1 into the Layer-1 v0.6.0 annotation schema, with per-clip narration included as language sidecars. 314,839 clips · 78,282,306 frames · 724.8 hours @ 30 fps · 129 tasks 100% narration coverage (1 sidecar per clip) 71 GB annotations + 2.3 GB narratives Videos are NOT included. This release contains annotations and narration only. Source video lives in griffinlabs/EgoDex-LeRobot-v3.0; orig_id in the manifest… See the full description on the dataset page: https://huggingface.co/datasets/VR-VLA/VR-egodex-annotation-converted-v6.0.tabularrobotics10M<n<100M0 likes427 downloads9d agoHugging Face10huyouare /SWE-bench_Verified_With_Annotationstabularn<1K1 likes426 downloads2y agoHugging Face11TyroneDragon /highlevel_thinking_with_grounding_annotation_split1000_v3_merged_promptstabular10K<n<100K0 likes374 downloads11mo agoHugging Face12windfromthenorth /scripted_atomic_train_frac_0.3_large_goal_annotationThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "ur5_wsg50_lego_atomic_step", "total_episodes": 664, "total_frames": 116214, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:664" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_train_frac_0.3_large_goal_annotation.tabularrobotics100K<n<1M0 likes334 downloads3mo agoHugging Face13houlab /cossmos-annotations-dbtabular1M<n<10M0 likes323 downloads1mo agoHugging Face14aiobservatory /annotations The AI Observatory A public measurement platform aggregating real-world AI conversations from seven sources under a unified 145-feature taxonomy. This dataset card accompanies paper The AI Observatory: A Public Measure of Real-World AI Use. 📄 Paper: [anonymous OpenReview link] 📊 Dashboard: https://project-ai-observatory.vercel.app/ 💾 Anonymous code: https://anonymous.4open.science/r/ai-observatory/README.md TL;DR 23,158 conversations, 85,633 turns, ~5,000… See the full description on the dataset page: https://huggingface.co/datasets/aiobservatory/annotations.tabulartext-classification100K<n<1M2 likes298 downloads1mo agoHugging Face15Synthyra /SwissProt-Annotation-Vocabulary Swiss-Prot Annotation Vocabulary 2026_02 This release converts a pinned Swiss-Prot snapshot into a versioned protein annotation vocabulary. Stable, namespaced term identifiers are the biological identity. Integer tokens are specific to this vocabulary and grammar version. Release summary Field Value Vocabulary version 2026_02-support10-v1 Grammar version 1 Swiss-Prot release 2026_02 Swiss-Prot release date 2026-06-10 Build date 2026-08-26… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/SwissProt-Annotation-Vocabulary.tabularfeature-extraction1M<n<10M1 likes262 downloads26d agoHugging Face16edesaras /CEFR-Sentence-Level-Annotations Dataset Card for Dataset Name 17k english sentences annotated by english education professionals. Original repo for CEFR-SP is located at this repo This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/edesaras/CEFR-Sentence-Level-Annotations.tabulartext-classification10K<n<100K6 likes247 downloads2y agoHugging Face17jasongraf1 /annotation_app_data Dataset Card for Systematic Review of Acceptability Judgments data A curated dataset of research articles used in a systematic review of judgment tasks in linguistics. Each entry records article-level metadata and experiment-level methodological features, supporting structured comparison and analysis across studies. Dataset Description This annotation dataset comprises systematically coded observations from a corpus of published studies employing judgment tasks in… See the full description on the dataset page: https://huggingface.co/datasets/jasongraf1/annotation_app_data.tabularn<1K0 likes203 downloads4d agoHugging Face18Alptekinege /TurkWeb-Edu-AnnotationsV3 TurkWeb-Edu V3 Model: Qwen/Qwen3-30B-A3B-Instruct-2507 Format: Structured JSON (vLLM 0.15.0) tabular100K<n<1M0 likes178 downloads6mo agoHugging Face19tvonarx /emboss-roof-annotations Emboss 3D Roof Reference Annotations Manual 3D reference meshes and editable annotations for Swiss and Brazilian buildings, prepared for the evaluation and parameter tuning of Emboss. The annotations describe building and roof geometry, including roof superstructures. Emboss source code 3dlabel annotation tool Emboss segmentation model Example reference annotation in 3dlabel: annotated mesh and LiDAR points (Figure D.1(a) in the paper). 3dn<1K0 likes175 downloads8d agoHugging Face20TyroneDragon /highlevel_thinking_with_grounding_annotation_split1000_v2_merged_promptstabular10K<n<100K0 likes168 downloads11mo agoHugging Face21nikhilchandak /gpqa-diamond-annotations GPQA Diamond Dataset This dataset contains filtered JSONL files of human annotations on question specificity, answer uniqueness, answer matching to the ground truth for different models for the GPQA Diamond dataset. The dataset was annotated by two human graders. It contains 198 (original size) * 2 = 396 rows as each rows is repeated twice (one for each human). A human grader given the question, actual answer and model response, has to answer whether the response matches the… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/gpqa-diamond-annotations.tabularn<1K1 likes148 downloads1y agoHugging Face22Synthyra /ESMC-6B-SAE-Annotation-Vocabulary-Features Vocabulary interpretations of ESMC-6B SAE features One row for every one of the 16,384 features of biohub/ESMC-6B-sae-layer60-k64-codebook16384, giving the protein annotation vocabulary term that best identifies what the feature detects, together with how well that identification holds on proteins the assignment never saw. This is the counterpart to biohub/ESMC-SAE-Features, produced without a language model. Where that release gives a free-text hypothesis per feature, this… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ESMC-6B-SAE-Annotation-Vocabulary-Features.tabularfeature-extraction100K<n<1M0 likes148 downloads26d agoHugging Face23Himpq /kwext-bilibili-video-title-annotations KwExt Bilibili Video Title Annotations This dataset is a model-assisted annotation set for the KwExt keyword extraction project. The current snapshot contains 5,000 Chinese Bilibili video titles from annotation stages video_title_zh_001 through video_title_zh_005, with 1,000 records in each stage. The release is intended for early experiments with: extracting title-grounded keywords and ranking their importance; broad semantic tags for retrieval and RAG metadata; dense tag… See the full description on the dataset page: https://huggingface.co/datasets/Himpq/kwext-bilibili-video-title-annotations.tabulartoken-classification1K<n<10K0 likes127 downloads5d agoHugging Face24michaljunczyk /bigos-annotationstabularn<1K1 likes122 downloads7mo agoHugging Face25ambient-intelligence-labs /egolongqa-synth-annotations EgoLongQA synthetic MCQs, teacher traces and annotation outputs Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026 EgoLongQA ≤2B track, other than the distillation set (which lives in infinitylogesh/egolongqa-junior-distill). ⚠️ Read this before counting rows The synthetic set is 943 questions over 408 videos, and it is stored two ways: file rows shape training_sets/train_synth_v3.jsonl 943 flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.tabularvisual-question-answering1K<n<10K0 likes121 downloads17d agoHugging Face26sed-i /mania-pattern-annotations osu!mania pattern annotations Snapshot v3 uses publication schema beatmap-lens-annotations version 4. This snapshot contains 592 human judgments, 4403 agent judgments, and 545 source identities (annotation and required calibration sources). Export implementation: GitHub commit 647009ab60ed. v3 release scope The default human layer contains the current effective human observations, with explicit High/Low confidence where recorded. The opt-in machine layer contains… See the full description on the dataset page: https://huggingface.co/datasets/sed-i/mania-pattern-annotations.tabular1K<n<10K1 likes115 downloads10d agoHugging Face27superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes113 downloads5mo agoHugging Face28LianeMarilin /4k-video-annotations 4K Video Annotations — Shot Segmentation and Camera Motion This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties. The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.imagen<1K0 likes110 downloads8d agoHugging Face29emarro /example_10kbp_human_annotationstabular100K<n<1M0 likes106 downloads1y agoHugging Face30YsK-dev /TurkWeb-Edu-AnnotationsV3 TurkWeb-Edu V3 Model: Qwen/Qwen3-30B-A3B-Instruct-2507 Format: Structured JSON (vLLM 0.15.0) tabular100K<n<1M0 likes104 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.