mylesmharrison/cornell-movie-dialog
Dataset Card for "cornell-movie-dialog" This is a reduced version of the Cornell Movie Dialog Corpus by Cristian Danescu-Niculescu-Mizil. The original dataset contains 220,579 conversational exchanges between 10,292 pairs of movie characters, involving 9,035 characters from 617 movies for a total 304,713 utterances. This reduced version of the dataset contains only the character tags and utterances from the movie_lines.txt file, with one utterance per line, suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/mylesmharrison/cornell-movie-dialog.
Dataset Card for "cornell-movie-dialog"
This is a reduced version of the Cornell Movie Dialog Corpus by Cristian Danescu-Niculescu-Mizil.
The original dataset contains 220,579 conversational exchanges between 10,292 pairs of movie characters, involving 9,035 characters from 617 movies for a total 304,713 utterances.
This reduced version of the dataset contains only the character tags and utterances from the movie_lines.txt file, with one utterance per line, suitable for training generative text models.
Dataset Description
- Homepage: https://www.cs.cornell.edu/~cristian/CornellMovie-DialogsCorpus.html
- Repository: https://convokit.cornell.edu/documentation/movie.html
- Paper: Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
