datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v3.0 Full (1,103 samples) - The complete DramaBench collection is now openly available, with context-continuation pairs designed to assess models across six independent evaluation dimensions.
Release Roadmap
Version
Samples
Status… See the full description on the dataset page: https://huggingface.co/datasets/FutureMa/DramaBench.DramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v2.0 (500 samples) - This release contains 500 carefully selected drama scripts with context-continuation pairs, designed to assess models across six independent evaluation dimensions. This represents a 5x expansion from v1.0, providing more comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/sxjiemust/DramaBench.DramaCV
Dataset Card for DramaCV
Dataset Summary
The DramaCV Dataset is an English-language dataset containing utterances of fictional characters in drama plays collected from Project Gutenberg. The dataset was automatically created by parsing 499 drama plays from the 15th to 20th century on Project Gutenberg, that are then parsed to attribute each character line to its speaker.
Task
This dataset was developed for Authorship Verification of literary characters. Each… See the full description on the dataset page: https://huggingface.co/datasets/gasmichel/DramaCV.DRAMA-X
Request access to the original DRAMA dataset athttps://usa.honda-ri.com/drama#Downloadthedataset
Download the ZIP and extract the fileintegrated_output_v2.json
Ensure you have your existingdrama_x_annotations.json (with empty image_path/video_path fields) in the same folder.
Run the population script:
python populate_drama_x.py \
drama_x_annotated.json \
integrated_output_v2.json \
-o drama_x_annotations_populated.json
DramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v1.0 (100 samples) - This is the initial release containing 100 carefully selected drama scripts with context-continuation pairs, designed to assess models across six independent evaluation dimensions.
Release Roadmap
Version
Samples
Status… See the full description on the dataset page: https://huggingface.co/datasets/tatan2/DramaBench.DramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v1.0 (100 samples) - This is the initial release containing 100 carefully selected drama scripts with context-continuation pairs, designed to assess models across six independent evaluation dimensions.
Release Roadmap
Version
Samples
Status… See the full description on the dataset page: https://huggingface.co/datasets/LizRob6913/DramaBench.meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Dramazy/meow-10k.beckett-dramatic-works-cpt
Samuel Beckett Complete Dramatic Works (CPT Dataset)
This dataset contains the complete, raw theatrical texts of Samuel Beckett's dramatic works.
It was specifically compiled for Continued Pre-Training (CPT) to teach Large Language Models the distinct vocabulary, pacing, and minimalist stage directions characteristic of Beckett's writing style.
Contents
This dataset consists of raw text blocks extracted from:
Waiting for Godot
Endgame
Krapp's Last Tape
Happy… See the full description on the dataset page: https://huggingface.co/datasets/alst10/beckett-dramatic-works-cpt.
