CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openeurollm /propella-annotations This dataset contains document annotations produced with propella-1-4b, a small multilingual LLM that annotates text documents across six categories: core content, classification, quality & value, audience & purpose, safety & compliance, and geographic relevance. The annotations can be used to filter, select, and curate LLM training data at scale. Properties Each document is annotated across 18 properties organized into six categories: Category Property Description… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/propella-annotations.text1B<n<10B20 likes7.4k downloads1mo agoHugging Face02kennethli319 /seamless-interaction-jefferson-annotations Seamless Interaction Jefferson-Style Annotations An automatic, turn-oriented annotation layer for the Meta Seamless Interaction Dataset. It compares the dataset's traditional transcript with an ASR-derived Jefferson-style condition and supplies speech-act, communicative-purpose, interactional-signal, alignment, and quality fields. This is a derived noncommercial research dataset. It does not redistribute the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.tabularautomatic-speech-recognition100K<n<1M0 likes873 downloads2mo agoHugging Face03HuggingFaceTB /python-edu-annotations Annotations for 📚 Python-Edu classifier This dataset contains the annotations used for training Python-Edu educational quality classifier. We prompt Llama-3-70B-Instruct to score python programs from StarCoderData based on their educational value. Note: the dataset contains the Python program, the prompt (using the first 1000 characters of the program) and the scores but it doesn't contain the full Llama 3 generation. text100K<n<1M2 likes816 downloads2y agoHugging Face04it-just-works /vast27m_annotations VAST-27M Annotations Dataset This dataset contains annotations from the VAST-27M dataset, originally created for the paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset" by Chen et al. (2024). Original Source This dataset is derived from the VAST-27M dataset, which was created by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. The original dataset and more… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/vast27m_annotations.tabular10M<n<100M1 likes741 downloads2y agoHugging Face05alakxender /od-syn-page-annotations-com 📦 Dhivehi Synthetic Document Layout + Textline Dataset This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script. Note: this version image are compressed. Raw version 📁 Repository: Hugging Face Datasets 📋 Dataset Summary Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.imageimage-classification10K<n<100K0 likes508 downloads1y agoHugging Face06HuggingFaceFW /fineweb-edu-llama3-annotations Annotations for 📚 FineWeb-Edu classifier This dataset contains the annotations used for training 📚 FineWeb-Edu educational quality classifier. We prompt Llama-3-70B-Instruct to score web pages from 🍷 FineWeb based on their educational value. Note: the dataset contains the FineWeb text sample, the prompt (using the first 1000 characters of the text sample) and the scores but it doesn't contain the full Llama 3 generation. text100K<n<1M50 likes470 downloads2y agoHugging Face07alakxender /od-syn-page-annotations 📦 Dhivehi Synthetic Document Layout + Textline Dataset This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis , visual document understanding , OCR fine-tuning, and related tasks specifically for Dhivehi script. 📋 Dataset Summary Total Examples: ~58,738 Image Content: Synthetic Dhivehi documents generated to simulate real-world layouts… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations.imagetext-classification10K<n<100K0 likes421 downloads1y agoHugging Face08mlfoundations-dev /merge_annotations_self_and_swegymtext100K<n<1M0 likes343 downloads1y agoHugging Face09houlab /cossmos-annotations-dbtabular1M<n<10M0 likes268 downloads1mo agoHugging Face10edesaras /CEFR-Sentence-Level-Annotations Dataset Card for Dataset Name 17k english sentences annotated by english education professionals. Original repo for CEFR-SP is located at this repo This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/edesaras/CEFR-Sentence-Level-Annotations.tabulartext-classification10K<n<100K6 likes243 downloads2y agoHugging Face11obalcells /longfact-annotationstext1K<n<10K2 likes231 downloads1y agoHugging Face12touati-kamel /forest-fire-annotations Forest Fire Detection Dataset — Auto-Annotated Bounding-box annotated version of touati-kamel/forest-fire-dataset, built for training forest-fire / smoke / fog object detection models. Overview This dataset contains video frames auto-labeled with bounding boxes for fire and smoke-related visual phenomena, using a zero-shot open-vocabulary object detector (Grounding DINO). It is derived from the original touati-kamel/forest-fire-dataset image classification dataset… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/forest-fire-annotations.image10K<n<100K0 likes228 downloads2mo agoHugging Face13ncbi /TrialGPT-Criterion-Annotationstext1K<n<10K7 likes225 downloads25d agoHugging Face14capitaletech /real-resumes-section-detection-annotationsimage1K<n<10K0 likes202 downloads8mo agoHugging Face15WIPI /deceptive_patterns_manual_annotationstext1K<n<10K0 likes195 downloads2y agoHugging Face16Alptekinege /TurkWeb-Edu-AnnotationsV3 TurkWeb-Edu V3 Model: Qwen/Qwen3-30B-A3B-Instruct-2507 Format: Structured JSON (vLLM 0.15.0) tabular100K<n<1M0 likes192 downloads6mo agoHugging Face17BowerApp /bower-waste-annotations Dataset Card for waste annotations made by the recycling solution Bower The data offered by Bower (Sugi Group AB) in collaboration with Google.org Dataset Summary The bower-waste-annotations dataset consists of 1440 images of waste and various consumer items taken by consumer phone cameras. The images are annotated with Material type and Object type classes, listed below. The images and annotations has been manually reviewed to ensure correctness. It is assumed… See the full description on the dataset page: https://huggingface.co/datasets/BowerApp/bower-waste-annotations.image1K<n<10K4 likes161 downloads2y agoHugging Face18H2KP /cdip-annotations-formnet-v1text100K<n<1M0 likes153 downloads4y agoHugging Face19owlgebra-ai /amz-image-annotationsimage1M<n<10M0 likes140 downloads8mo agoHugging Face20obalcells /longfact-augmented-annotationstext10K<n<100K0 likes137 downloads1y agoHugging Face21hfmlsoc /sp500_earnings_annotationstextn<1K0 likes131 downloads9mo agoHugging Face22HuggingFaceFW /ocr-annotations PDF OCR Classification Dataset This dataset contains PDF documents with annotations for OCR classification tasks. Dataset Structure Each row contains: filename: Original PDF filename pdf: PDF file as binary data (using Pdf feature type) class: Binary classification label (OCR/NOCR) truncation_type: Whether the PDF is truncated or non-truncated pdf_size_bytes: Size of the PDF file in bytes Class Distribution class NOCR 1393 OCR 227 Usage… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/ocr-annotations.document1K<n<10K19 likes129 downloads11mo agoHugging Face23hfmlsoc /sp500_sec_10k_annotationstextn<1K0 likes122 downloads9mo agoHugging Face24michaljunczyk /bigos-annotationstabularn<1K1 likes121 downloads7mo agoHugging Face25sed-i /mania-pattern-annotations osu!mania pattern annotations Snapshot v3 uses publication schema beatmap-lens-annotations version 4. This snapshot contains 592 human judgments, 4403 agent judgments, and 545 source identities (annotation and required calibration sources). Export implementation: GitHub commit 647009ab60ed. v3 release scope The default human layer contains the current effective human observations, with explicit High/Low confidence where recorded. The opt-in machine layer contains… See the full description on the dataset page: https://huggingface.co/datasets/sed-i/mania-pattern-annotations.tabular1K<n<10K2 likes117 downloads12d agoHugging Face26tvosch /GPT-NL-propella-annotations GPT-NL Propella Annotations Document-level quality and content annotations for the Dutch subset of GPT-NL/GPT-NL_Public_Corpus. What is Propella? Propella (ellamind/propella-1-4b) is a 4B-parameter language model fine-tuned from Qwen3 for document-level annotation of LLM pretraining data. Given a document in any language, it produces a structured JSON object with 18 quality, classification, and safety properties (see below for the features). These annotations can… See the full description on the dataset page: https://huggingface.co/datasets/tvosch/GPT-NL-propella-annotations.text10M<n<100M4 likes114 downloads17d agoHugging Face27emarro /example_10kbp_human_annotationstabular100K<n<1M0 likes105 downloads1y agoHugging Face28ShantyCam /objects365-annotations Dataset Card for "objects365" More Information needed text1M<n<10M0 likes102 downloads2mo agoHugging Face29YsK-dev /TurkWeb-Edu-AnnotationsV3 TurkWeb-Edu V3 Model: Qwen/Qwen3-30B-A3B-Instruct-2507 Format: Structured JSON (vLLM 0.15.0) tabular100K<n<1M0 likes101 downloads7mo agoHugging Face30north /scandinavian-educational-annotations Scandinavian Educational Annotations Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash. texttext-generation100K<n<1M3 likes100 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.