datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VADBThe data from our video dataset is stored in VADB.zip
Our train.csv contains all annotated data related to aesthetic scores.
For text annotations such as comments and tags mentioned in the paper, please refer to merged_comment_tag.json.
[2025/10/22 UPDATE] The annotation files are fully open-sourced, but due to copyright issues with some videos, 2,609 video clips have been retained and not open-sourced, while the remaining 7,881 video clips have been fully released.
skyrimdialogstestHebrew_VAD_lexicon
Hebrew VAD Lexicon
The Hebrew VAD Lexicon is an enhanced version of an automatically translated affective lexicon, originally derived from the English VAD lexicon created by Mohammad (2018) .
It provides valence, arousal, and dominance (VAD) scores for Hebrew words.
The lexicon was carefully curated by manually reviewing and correcting the automatic translations and enriching the dataset with additional linguistic information.
There are two versions of the dataset:
English-Hebrew… See the full description on the dataset page: https://huggingface.co/datasets/GiliGold/Hebrew_VAD_lexicon.FER2013-VAD-annotationThis dataset involves train-20240123-14902.csv, publictest-20240508.csv and privatetest-20240506-yh.csv, which could be used for public and private test respectively. 14902, 1298 and 3589 samples are for train, public and private dataset at present.
journal-entries-emotion-detection-vadReddit Diary of a Redditor VAD Dataset
Dataset Creation Process
Scraping Reddit Posts
Posts were scraped from the r/diaryofaredditor subreddit using the Reddit API.
The script used for scraping is shown below:import requests
import csv
import time
access_token = ""
headers = {
"Authorization": f"bearer {access_token}",
"User-Agent": "ChangeMeClient/0.1"
}
url = "https://oauth.reddit.com/r/diaryofaredditor/new"
params = {"limit": 100}
after = None
csv_path =… See the full description on the dataset page: https://huggingface.co/datasets/mmarkusmalone/journal-entries-emotion-detection-vad.jigsaw-toxic-comment-vad-zh
Jigsaw Toxic Comment VAD (EN-ZH)
Dataset Summary
This dataset is derived from the Kaggle competition "Jigsaw Multilingual Toxic Comment Classification".
It contains English comments, adjusted VAD scores (valence/arousal/dominance) for each comment, and Chinese translations.
Translation failures are provided separately.
Languages
English (en) and Chinese (zh).
Data Files
train.csv: 221,858 rows with successful translations.
failed.csv: 1,691 rows… See the full description on the dataset page: https://huggingface.co/datasets/Pectics/jigsaw-toxic-comment-vad-zh.dialogsum
Dataset Card for DIALOGSum Corpus
Dataset Description
Links
Homepage: https://aclanthology.org/2021.findings-acl.449
Repository: https://github.com/cylnlp/dialogsum
Paper: https://aclanthology.org/2021.findings-acl.449
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/vadrami/dialogsum.cognitive-matrices-observationsSCIENTIFIC AND TECHNICAL DATA ABSTRACT
Exclusive Chronological Dataset of Cognitive Matrices and Long-Term Systemic Observations (2010–2026)
Data Specification and Validation
Asset Type: A continuous 15-year time-series dataset consisting of structured empirical records (141 pages of dense text in Microsoft Word format).
Chronological Accuracy: The dataset is maintained in a strict sequential order. Each control point is logged with the exact month, year, and hour of… See the full description on the dataset page: https://huggingface.co/datasets/Vadim067/cognitive-matrices-observations.lezgi-books-russian-parallel
Lezgi Books Lezgi-Russian Parallel Corpus
Dataset Summary
This corpus was made by translating Lezgi books texts into Russian and aggregating the resulting parallel TSV files.
The published dataset is packaged as a strict 2-column TSV for Hugging Face compatibility.
Source language: Lezgi
Target language: Russian
Format: TSV
Columns: lezgi, russian
Dataset Structure
Files:
lezgi-books-russian-parallel.tsv
Columns:
lezgi
russian
Size
Total rows… See the full description on the dataset page: https://huggingface.co/datasets/vadim-pashaev/lezgi-books-russian-parallel.AudioProjAnthropicInterviewer
Anthropic Interviewer
A tool for conducting AI-powered qualitative research interviews at scale. In this study, we used Anthropic Interviewer to explore how 1,250 professionals integrate AI into their work and how they feel about its role in their future.
Associated Research: Introducing Anthropic Interviewer: What 1,250 professionals told us about working with AI
Dataset
This repository contains interview transcripts from 1,250 professionals:
General Workforce (N=1… See the full description on the dataset page: https://huggingface.co/datasets/Vadiyala/AnthropicInterviewer.sentinel-nids-telemetrylezgichal-lezgi-russian-parallel
LezgiChal.ru Lezgi-Russian Parallel Corpus
Dataset Summary
This dataset contains Lezgi texts from lezgichal.ru paired with Russian translations in TSV format.
Source language: Lezgi
Target language: Russian
Format: TSV
Columns: lezgi, russian
Dataset Structure
The dataset is distributed as a single TSV file:
lezgichal-lezgi-russian-parallel.tsv
The file has two columns:
lezgi
russian
Creation Notes
The source material was collected from… See the full description on the dataset page: https://huggingface.co/datasets/vadim-pashaev/lezgichal-lezgi-russian-parallel.
