datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DreamBank-annotated
Presentation
DreamBank, an open corpus of more than 27,000 dream narratives, mostly written in English.
Annotations were produced using dream-t5, a LaMini-Flan-T5 model finetuned on Hall and Van de Castle annotations to predict character and emotion. I've introduced this task in this paper:
Gustave Cortal. 2024. Sequence-to-Sequence Language Models for Character and Emotion Detection in Dream Narratives. In Proceedings of the 2024 Joint International Conference on Computational… See the full description on the dataset page: https://huggingface.co/datasets/gustavecortal/DreamBank-annotated.green_claims_annotatedovenbird-annotated-1k
1000 3-second audio clips annotated for Ovenbird song presence
This dataset contains 100 3 second audio clips randomly selected from a 4-year passive acoustic monitoring dataset at 126 recording location in Pennsylvania. Experts familiar with Ovenbird song reviewed the audio and spectrograms using the Dipper app. The clip_annotations.csv file contains the annotation column with 'yes' for confirmed presence (N=92), 'no' for confirmed absence (894), or 'uncertain' if presence of… See the full description on the dataset page: https://huggingface.co/datasets/sammlapp/ovenbird-annotated-1k.annotatebench-results
AnnotateBench Results
AnnotateBench is a cost-aware benchmark for comparing annotation strategies
across public text-classification datasets and label budgets. This repository
contains derived experiment results and aggregate statistical summaries. It
does not redistribute the source datasets.
Before making this repository public: replace this notice after the Anote
AI team has selected a release license and verified that the derived results
may be distributed under that… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/annotatebench-results.English_Malayalam_Translation_Human_annotated
English-Malayalam Government Parallel Corpus Synth
This dataset contains synthetic machine-translated English-Malayalam sentence pairs aligned from government and administrative text.
Machine Translation Notice
All parallel text in this dataset should be treated as synthetic machine-translated data. It is intended for research, corpus filtering, model adaptation, and experimentation. It should not be treated as human-verified gold translation without additional… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/English_Malayalam_Translation_Human_annotated.example_annotated_code_repo_dataA description of the fields:
Column
What it captures
Typical values
id
Row identifier
1-100
repo_name
Example repository label
repo_14
file_path
Path + filename with extension
src/utils/parsefile.py
language
Programming language
Python, Java…
function_name
Target symbol that was reviewed
validateSession
annotation_summary
Free-text note written by the annotator
“Added input validation…”
potential_bug
Did the annotator flag a likely bug? (Yes/No)… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/example_annotated_code_repo_data.imagenet_safety_annotatedThis is a safety annotation set for ImageNet. It uses the LlavaGuard-13B model for annotating.
The annotations entail a safety category (image-category), an explanation (assessment), and a safety rating (decision). Furthermore, it contains the unique ImageNet id class_sampleId, i.e. n04542943_1754.
These annotations allow you to train your model on only safety-aligned data. Plus, you can define yourself what safety-aligned means, i.e. discard all images where decision=="Review Needed" or… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/imagenet_safety_annotated.annotated_news_summaryThis dataset is created for instruction tuning purpose.It is based on the News Summarization dataset.
The instructions are given in the inputs column and their completions/answers are provided in the targets column. The template_id tracks each input_template-target_template pair. There are 15 template ids (from 1 to 15).
The ID and their respective templates are given below. no_template indicates that no template was used and only the summary or direct answer was provided for that input.… See the full description on the dataset page: https://huggingface.co/datasets/TahmidH/annotated_news_summary.COVID-19-conspiracy-theories-annotated-tweets
Dataset Card for Dataset Name
The dataset includes approximately 54,000 tweet IDs collected through the Twitter API between November 2019 and December 2021, along with annotations indicating whether a tweet supports a particular conspiracy theory.
Dataset Details
Curated by: Izabela Krysinska
Funded by [optional]: EEA Financial Mechanism 2014–2021. Project registration number: 2019/35/J/HS6/03498
Shared by [optional]: Izabela Krysinska
Language(s) (NLP): English… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-conspiracy-theories-annotated-tweets.Group_M_An_Annotated_Dataset_for_Hearing-Impaired_Speech-to-Text_Correction
Title
An Annotated Dataset for Hearing-Impaired Speech-to-Text Correction
Abstract
This dataset is specifically designed for the task of correcting speech to text errors in hearing-impaired individuals.
It includes one hour of real speech files of hearing-impaired individuals, automatic speech recognition (ASR) output text, and manually corrected standard text.
We searched for an hour of audio from hearing-impaired individuals to ensure that the voice was authentic and… See the full description on the dataset page: https://huggingface.co/datasets/llllliuuy/Group_M_An_Annotated_Dataset_for_Hearing-Impaired_Speech-to-Text_Correction.formatted_annotated_addiction_counseling_csv_SFTSyntactic-Semantic-Annotated-Italian-Corpus
Annotazione Sintattico-Funzionale e Disambiguazione della Lingua Italiana
This dataset was generated by fetching random first paragraphs from Italian Wikipedia (it.wikipedia.org)
and then processing them using Gemini AI with the following goal:
Processing Goal: riduci la ambiguità aggiungi tag grammaticali (soggetto) (verbo) eccetera. e tag funzionali es. (indica dove è nato il soggetto) (indica che il soggetto possiede l'oggetto) eccetera
Source Language: Italian (from Wikipedia)… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/Syntactic-Semantic-Annotated-Italian-Corpus.Annotated_Food_Vlog_Dataset_GroupL
Dataest Description
This project has constructed a multimodal corpus of language strategies for food exploration videos on Chinese social media. The dataset is centered around the videos of the well-known blogger "Diao Yueshe Shi Yu Ji", containing approximately 1,000 entries with a total of 90 minutes of transcribed video content. The dataset is stored in CSV format and meticulously records the original dialogue, synthetic text generated by large language models (LLMs), rhetorical… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/Annotated_Food_Vlog_Dataset_GroupL.flickr8k-sau-pace-annotated
Annotation
Annotated this dataset by clasifying the images into slow, medium or fast depending on the suitable paced background music.
annotated_training_antisemannotated training data for text classification on antisemitism
sci-sentiment-annotated-datasetThis dataset contains manual sentiment annotations within a scientific context, with labels "p," "n," and "o" denoting positive, negative, and others, respectively. The dataset comprises 100 records and was created for evaluating the performance of the sci-sentiment-classify model : https://huggingface.co/puzzz21/sci-sentiment-classify.
annotated_switchboard_v1alpaca-cleaned-annotatedr1_annotated_math-mistralgender-annotated-tokensNLI-SK-annotated
Dataset Card for NLI-SK-annotated
The dataset consists of 800 premise-hypothesis pairs that were handwritten and annotated by two people, excluding the author, for the task of natural language inference in Slovak.
The goal of creating this dataset is to provide a set of labeled and annotated human data that can be used to test abilities of fine-tuned language models. To our knowledge, other NLI datasets in Slovak are translated.
The dataset contains 204 entailment pairs, 215… See the full description on the dataset page: https://huggingface.co/datasets/natalia-nk/NLI-SK-annotated.annotated-math-tutoring-datasethuman_annotated_datachatbot_arena_harmful_annotatedThis dataset is derived from lmsys' Chatbot Arena Conversations Dataset.
This dataset contains the following columns:
question_id: The same column from the original dataset, to facilitate joining with the original dataset
user_prompt: The first user message from the original dataset's conversation_a column
is_harmful: Binary label indicating whether user_prompt contains harmful intent or not
harmful_judge_model: The LLM used as judge to generate the label is_harmful
Some rows (135) weren't… See the full description on the dataset page: https://huggingface.co/datasets/felipeadachi/chatbot_arena_harmful_annotated.annotated-mimic-cxr-reportsExplainLikeIm5_Annotated_Dataset
LLM Argumentation Preference Dataset
Dataset created for the NLP Research Course 097920 (Technion).
Each example includes a user query and two responses annotated by 3 human annotators for preference, source identification etc.
🧩 Tasks
The dataset includes four main annotation tasks:
Preference Task – Which response is easier to understand?
Source Identification Task – Which response is written by a human or an AI?
Appeal to Expert Task – Does the response's writer… See the full description on the dataset page: https://huggingface.co/datasets/SlowSenik/ExplainLikeIm5_Annotated_Dataset.ner-annotated-linux-product-issue-title-description.csvmetra_western_european_drama_annotated_corporaworkplace-emotion-annotated-v1
NLP Emotional Drift & Ethical Framing Evaluation
This repository contains a manually annotated set of model-generated responses to open-ended prompts. The goal is to evaluate subtle issues in AI-generated answers such as emotional misalignment, ethical distortion, responsibility evasion, and justification drift.
Contents
prompts_and_responses.csv: 10 original prompts and ChatGPT baseline answers
annotations.csv: Manual evaluations with binary flags and rationale… See the full description on the dataset page: https://huggingface.co/datasets/Saranarunkumarak/workplace-emotion-annotated-v1.
