datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
propella-annotations
This dataset contains document annotations produced with propella-1-4b, a small multilingual LLM that annotates text documents across six categories: core content, classification, quality & value, audience & purpose, safety & compliance, and geographic relevance. The annotations can be used to filter, select, and curate LLM training data at scale.
Properties
Each document is annotated across 18 properties organized into six categories:
Category
Property
Description… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/propella-annotations.seamless-interaction-jefferson-annotations
Seamless Interaction Jefferson-Style Annotations
An automatic, turn-oriented annotation layer for the
Meta Seamless Interaction Dataset.
It compares the dataset's traditional transcript with an ASR-derived
Jefferson-style condition and supplies speech-act, communicative-purpose,
interactional-signal, alignment, and quality fields.
This is a derived noncommercial research dataset. It does not redistribute
the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.python-edu-annotations
Annotations for 📚 Python-Edu classifier
This dataset contains the annotations used for training Python-Edu educational quality classifier. We prompt Llama-3-70B-Instruct to score python programs from StarCoderData based on their educational value.
Note: the dataset contains the Python program, the prompt (using the first 1000 characters of the program) and the scores but it doesn't contain the full Llama 3 generation.
vast27m_annotations
VAST-27M Annotations Dataset
This dataset contains annotations from the VAST-27M dataset, originally created for the paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset" by Chen et al. (2024).
Original Source
This dataset is derived from the VAST-27M dataset, which was created by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. The original dataset and more… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/vast27m_annotations.od-syn-page-annotations-com
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script.
Note: this version image are compressed.
Raw version 📁 Repository: Hugging Face Datasets
📋 Dataset Summary
Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.fineweb-edu-llama3-annotations
Annotations for 📚 FineWeb-Edu classifier
This dataset contains the annotations used for training 📚 FineWeb-Edu educational quality classifier. We prompt Llama-3-70B-Instruct to score web pages from 🍷 FineWeb based on their educational value.
Note: the dataset contains the FineWeb text sample, the prompt (using the first 1000 characters of the text sample) and the scores but it doesn't contain the full Llama 3 generation.
od-syn-page-annotations
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis , visual document understanding , OCR fine-tuning, and related tasks specifically for Dhivehi script.
📋 Dataset Summary
Total Examples: ~58,738
Image Content: Synthetic Dhivehi documents generated to simulate real-world layouts… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations.merge_annotations_self_and_swegymcossmos-annotations-dbCEFR-Sentence-Level-Annotations
Dataset Card for Dataset Name
17k english sentences annotated by english education professionals. Original repo for CEFR-SP is located at this repo
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/edesaras/CEFR-Sentence-Level-Annotations.longfact-annotationsforest-fire-annotations
Forest Fire Detection Dataset — Auto-Annotated
Bounding-box annotated version of touati-kamel/forest-fire-dataset,
built for training forest-fire / smoke / fog object detection models.
Overview
This dataset contains video frames auto-labeled with bounding boxes for fire and
smoke-related visual phenomena, using a zero-shot open-vocabulary object detector
(Grounding DINO). It is derived from the original touati-kamel/forest-fire-dataset image
classification dataset… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/forest-fire-annotations.TrialGPT-Criterion-Annotationsreal-resumes-section-detection-annotationsdeceptive_patterns_manual_annotationsTurkWeb-Edu-AnnotationsV3
TurkWeb-Edu V3
Model: Qwen/Qwen3-30B-A3B-Instruct-2507
Format: Structured JSON (vLLM 0.15.0)
bower-waste-annotations
Dataset Card for waste annotations made by the recycling solution Bower
The data offered by Bower (Sugi Group AB) in collaboration with Google.org
Dataset Summary
The bower-waste-annotations dataset consists of 1440 images of waste and various consumer items taken by consumer phone cameras. The images are annotated with Material type and Object type classes, listed below.
The images and annotations has been manually reviewed to ensure correctness. It is assumed… See the full description on the dataset page: https://huggingface.co/datasets/BowerApp/bower-waste-annotations.cdip-annotations-formnet-v1amz-image-annotationslongfact-augmented-annotationssp500_earnings_annotationsocr-annotations
PDF OCR Classification Dataset
This dataset contains PDF documents with annotations for OCR classification tasks.
Dataset Structure
Each row contains:
filename: Original PDF filename
pdf: PDF file as binary data (using Pdf feature type)
class: Binary classification label (OCR/NOCR)
truncation_type: Whether the PDF is truncated or non-truncated
pdf_size_bytes: Size of the PDF file in bytes
Class Distribution
class
NOCR 1393
OCR 227
Usage… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/ocr-annotations.sp500_sec_10k_annotationsbigos-annotationsmania-pattern-annotations
osu!mania pattern annotations
Snapshot v3 uses publication schema beatmap-lens-annotations version 4.
This snapshot contains 592 human judgments, 4403 agent judgments, and 545 source identities (annotation and required calibration sources).
Export implementation: GitHub commit 647009ab60ed.
v3 release scope
The default human layer contains the current effective human observations, with
explicit High/Low confidence where recorded. The opt-in machine layer contains… See the full description on the dataset page: https://huggingface.co/datasets/sed-i/mania-pattern-annotations.GPT-NL-propella-annotations
GPT-NL Propella Annotations
Document-level quality and content annotations for the Dutch subset of GPT-NL/GPT-NL_Public_Corpus.
What is Propella?
Propella (ellamind/propella-1-4b) is a 4B-parameter language model fine-tuned from Qwen3 for document-level annotation of LLM pretraining data. Given a document in any language, it produces a structured JSON object with 18 quality, classification, and safety properties (see below for the features). These annotations can… See the full description on the dataset page: https://huggingface.co/datasets/tvosch/GPT-NL-propella-annotations.example_10kbp_human_annotationsobjects365-annotations
Dataset Card for "objects365"
More Information needed
TurkWeb-Edu-AnnotationsV3
TurkWeb-Edu V3
Model: Qwen/Qwen3-30B-A3B-Instruct-2507
Format: Structured JSON (vLLM 0.15.0)
scandinavian-educational-annotations
Scandinavian Educational Annotations
Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash.
