CoolFace
3 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01antfr99 /hitchcock-psycho-1960-film-dataset-transformed Psycho → AI-Model Dataset (Transformed) A thematic re-skin of the Psycho (1960) Q&A dataset into an original AI-model setting where the world is transformed into an AI/data-center environment. Character names, actor names, objects, locations, production references, dates, and thematic elements are remapped to AI/ML concepts and modern technology. File: psycho_dataset_transformed.jsonl Format: JSONL — one JSON object per line Schema: each line has prompt and completion string… See the full description on the dataset page: https://huggingface.co/datasets/antfr99/hitchcock-psycho-1960-film-dataset-transformed.texttext-generation1K<n<10K0 likes141 downloads10d agoHugging Face02BAAI /IndustryCorpus_film[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_film.texttext-generation10M<n<100M3 likes109 downloads1mo agoHugging Face03kazkiryuu /Films Screenplay Corpus — Tokenized (GPT-2) Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required. Dataset Description This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/kazkiryuu/Films.text-generation100K<n<1M1 likes29 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.