CoolFace
Datasetpublic

kazkiryuu/Films

Screenplay Corpus — Tokenized (GPT-2) Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required. Dataset Description This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/kazkiryuu/Films.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
1likes29downloads
1 commits on main
6f617bd4mo ago

Duplicate from raghavnimbalkar/movie-screenplays-tokenized-dataset

kazkiryuu, raghavnimbalkar