datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
screenplay-features
Screenplay Scene Salience Features
Pre-extracted linguistic and narrative features for screenplay scene salience detection from the MENSA dataset.
Dataset Description
This dataset contains 913 linguistic features extracted from movie screenplays in the MENSA dataset. Features are organized into 24 feature groups covering various aspects of linguistic, narrative, and discourse analysis.
Dataset Statistics
Split
Samples
Size
Train
117,503
172.9 MB… See the full description on the dataset page: https://huggingface.co/datasets/Ishaank18/screenplay-features.screenplay-features-linguistic
Screenplay Features - Linguistic Categories
This dataset reorganizes the features from screenplay-features into theoretically-motivated linguistic categories.
Dataset Structure
The dataset contains 837 features organized into 10 linguistic categories:
1. SURPRISAL (57 features)
Language model predictability features measuring cognitive processing difficulty.
bert_surprisal (15)
surprisal (5) - GPT-2 surprisal
gpt2_char_surprisal (6)
ngram_surprisal (5)… See the full description on the dataset page: https://huggingface.co/datasets/Ishaank18/screenplay-features-linguistic.imsdb-drama-screenplayscreen_play_for_personaimsdb-horror-screenplaybuzz_sources_084_processed_unified_plot_screenplay_books_dialogimsdb-comedy-screenplayimsdb-sci_fi-screenplay
