datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
movies
Movie Scripts Dataset
The Movie Scripts Dataset consists of scripts from 1,172 movies, providing a comprehensive collection of movie dialogues and narratives. This dataset is designed to support various natural language processing (NLP) tasks, including dialogue generation, script summarization, and text analysis.
Details
The dataset contains 2 columns:
Name: The title of the movie.
Script: The full script of the movie in English.
Usage
The Movie Scripts… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/movies.movie-screenplays-tokenized-dataset
Screenplay Corpus — Tokenized (GPT-2)
Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required.
Dataset Description
This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/raghavnimbalkar/movie-screenplays-tokenized-dataset.Haruhi-Zero-RolePlaying-movie-PIPPA
2000 Chinese RoleCards from IMDB_250 Movies and PIPPA
用于拓展zero-shot角色扮演的角色卡片。
其中870个角色来自电影字幕总结(id为movie_xx),其中406张翻译成了简体中文,剩下的没翻(所以有些繁体或者英文混杂)
1270个角色来自于对PIPPA数据集的翻译
凌云志@伯恩茅斯大学 使用射手api爬取了电影的字幕
李鲁鲁 完成了从字幕到角色卡片的总结,以及对数据的翻译(openai)
后续
我们后续打算用这些卡片 从openai, CharacterGLM, KoboldAI的api中,利用Baize的方式去获得数据。
项目主页 https://github.com/LC1332/Chat-Haruhi-Suzumiya
如果你要讨论加入我们的项目
可以把你的联系方式私信发给 https://www.zhihu.com/people/cheng-li-47
movies
Movie Scripts Dataset
The Movie Scripts Dataset consists of scripts from 1,172 movies, providing a comprehensive collection of movie dialogues and narratives. This dataset is designed to support various natural language processing (NLP) tasks, including dialogue generation, script summarization, and text analysis.
Details
The dataset contains 2 columns:
Name: The title of the movie.
Script: The full script of the movie in English.
Usage
The Movie Scripts… See the full description on the dataset page: https://huggingface.co/datasets/chuanli1013/movies.letterboxd-all-movie-data
Letterboxd Film Dataset
This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings.
Dataset Summary
Total Films: 847,209
File Size: ~1.12 GB (1,120,572,122 bytes)
Format: JSONL (JSON Lines)
Language: Primarily English, with some multilingual content
Data Structure
Each line contains a JSON object with the following fields:
{
"url":… See the full description on the dataset page: https://huggingface.co/datasets/PratikDhonde/letterboxd-all-movie-data.MovieStoryGen
MovieStoryGen: Movie-Inspired Creative Writing Dataset
Dataset Description
MovieStoryGen is a high-quality dataset for evaluating and fine-tuning large language models on creative story generation. The dataset contains structured writing prompts paired with detailed story responses that are inspired by IMDb's top 250 movies. Each entry contains a creative writing prompt and a corresponding well-crafted story that reimagines the essence of a classic film in a new context.… See the full description on the dataset page: https://huggingface.co/datasets/FutureMa/MovieStoryGen.turkish-comprehensive-movie-series-dataset
Beyazperde Film & Series Dataset
This dataset contains a comprehensive collection of Turkish films and TV series from Beyazperde.com, including detailed information about movies, series, cast, reviews, and ratings.
Dataset Summary
Total Movies: 27,227
Total Series: 11,240
Total Entries: 38,467
File Size: ~222 MB
Format: JSONL (JSON Lines)
Language: Turkish
Source: Beyazperde.com
Data Structure
Each line in the JSONL file contains a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/turkish-comprehensive-movie-series-dataset.50k_IMBD_Movie_Review_by_HNM
🎬 50K IMDB Movie Reviews Dataset
A balanced sentiment analysis dataset with 50,000 IMDB movie reviews
Overview • Dataset Structure • Statistics • Usage • Citation
📖 Overview
This dataset contains 50,000 movie reviews from IMDB, perfectly balanced between positive and negative sentiments. Each review includes the original text, reviewer rating (1-10), sentiment label, and source URL, making it ideal for sentiment analysis, text classification, and NLP research.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/50k_IMBD_Movie_Review_by_HNM.movies
Movie Scripts Dataset
The Movie Scripts Dataset consists of scripts from 1,172 movies, providing a comprehensive collection of movie dialogues and narratives. This dataset is designed to support various natural language processing (NLP) tasks, including dialogue generation, script summarization, and text analysis.
Details
The dataset contains 2 columns:
Name: The title of the movie.
Script: The full script of the movie in English.
Usage
The Movie Scripts… See the full description on the dataset page: https://huggingface.co/datasets/RobbieHurst/movies.malayalam_movies_sample
Malayalam Movies Instruction Dataset 🎬
This dataset contains a small set of Malayalam movie–related Q&A instructions.It can be used for fine-tuning instruction-based models such as Llama, Mistral, or GPT-like systems.
📂 Dataset Structure
Each record has three fields:
Field
Description
instruction
The prompt or question.
input
Optional context or input text (may be empty).
output
The expected response.
🧾 Example
{
"instruction": "Give… See the full description on the dataset page: https://huggingface.co/datasets/Brinda12/malayalam_movies_sample.
