CoolFace
Datasetpublic

raghavnimbalkar/movie-screenplays-tokenized-dataset

Screenplay Corpus — Tokenized (GPT-2) Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required. Dataset Description This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/raghavnimbalkar/movie-screenplays-tokenized-dataset.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
2likes50downloads
7 commits on main
8fcb4594mo ago

Update README.md

raghavnimbalkar
7cddda94mo ago

Update README.md

raghavnimbalkar
1e9d25e4mo ago

Update README.md

raghavnimbalkar
40ccdc64mo ago

Update README.md

raghavnimbalkar
a81937a4mo ago

Create README.md

raghavnimbalkar
d8d91744mo ago

Upload 4 files

raghavnimbalkar
67d64664mo ago

initial commit

raghavnimbalkar