CoolFace
Datasetpublic

raghavnimbalkar/movie-screenplays-tokenized-dataset

Screenplay Corpus — Tokenized (GPT-2) Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required. Dataset Description This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/raghavnimbalkar/movie-screenplays-tokenized-dataset.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
2likes50downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
raghavnimbalkar/movie-screenplays-tokenized-dataset · CoolFace