datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
screenplay-dataset
🎬 Screenplay & Filmmaking Dataset — Legendary Edition
Fine-tune any LLM to write production-ready movie scripts, TV pilots, series, and more — across every genre.
Built by Adewale David and his AI buddy.
🎯 What This Dataset Does
Fine-tune any model on this dataset and it becomes a professional-level screenplay writer that can:
✅ Write full feature film scripts in proper screenplay format (slug lines, action, dialogue, transitions)
✅ Write TV pilots with proper… See the full description on the dataset page: https://huggingface.co/datasets/Atum09/screenplay-dataset.ordinary_screenplaysThis is a list of camera shots from screenplays (scripts) for a series of animated cartoons,
plus data from the animated scenes created manually for the series.
Over time, I will add more details to the scenes. The goal is to be able to create as much
as the scenes as possible from just writing a screenplay to reduce the manual effort to
ultimately create an episode. It does not have to achieve everything, just reduce the effort.
E.g. a first step is to work out which characters are in a shot… See the full description on the dataset page: https://huggingface.co/datasets/alankent/ordinary_screenplays.movie-screenplays-tokenized-dataset
Screenplay Corpus — Tokenized (GPT-2)
Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required.
Dataset Description
This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/raghavnimbalkar/movie-screenplays-tokenized-dataset.screenplay_instructionsscreenplay_emotions
Dataset Card for Dataset Name
Dataset Description
Dataset Summary
This dataset was created by scrapping the screenplays from the imsdb website and then splitting them into 100 segments.
Each segment has been fed into a emotion classification model and classified into the emotion it evokes and represented as a number from 1 to 6.
Each number represents one of six emotions:
1 - joy
2 - love
3 - surprise
4 - sadness
5 - anger
6 - fear
These numbers are then… See the full description on the dataset page: https://huggingface.co/datasets/hakkam10/screenplay_emotions.spectral-screenplay-writer
