datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
screenplay-dataset
🎬 Screenplay & Filmmaking Dataset — Legendary Edition
Fine-tune any LLM to write production-ready movie scripts, TV pilots, series, and more — across every genre.
Built by Adewale David and his AI buddy.
🎯 What This Dataset Does
Fine-tune any model on this dataset and it becomes a professional-level screenplay writer that can:
✅ Write full feature film scripts in proper screenplay format (slug lines, action, dialogue, transitions)
✅ Write TV pilots with proper… See the full description on the dataset page: https://huggingface.co/datasets/Atum09/screenplay-dataset.movie-screenplays-tokenized-dataset
Screenplay Corpus — Tokenized (GPT-2)
Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required.
Dataset Description
This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/raghavnimbalkar/movie-screenplays-tokenized-dataset.
