datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DramaADDramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v3.0 Full (1,103 samples) - The complete DramaBench collection is now openly available, with context-continuation pairs designed to assess models across six independent evaluation dimensions.
Release Roadmap
Version
Samples
Status… See the full description on the dataset page: https://huggingface.co/datasets/FutureMa/DramaBench.ai-drama-production-harness-demo-media
The Second Key — demo media
This local staging tree contains reused AI-generated reference art and
model-rendered video from the operator-owned project run.
QC disclosure
The operator explicitly skipped measured per-clip QC for this set
(projects/default/runs/qc_skipped.json, schema version 1,
reason operator_satisfied). This is an operator QC waiver, not a fabricated
passing QC verdict. The staging gate independently resolved the current render
source and… See the full description on the dataset page: https://huggingface.co/datasets/tungmtp/ai-drama-production-harness-demo-media.Hongguo-Short-Drama-Corpus-AI-Labeled
🎬 2025年红果短剧全量语料库 (AI标注版 V1)
Hongguo Short Drama Corpus with AI-Augmented Audience Labels
1. 数据集简介 (Dataset Summary)
本数据集包含了约 1500 条来自红果 (Hongguo/RedFruit) 平台的微短剧精选数据。本数据集旨在为中文短剧的 NLP 研究、市场趋势分析以及自动化剧名生成等任务提供高质量的基准数据。
核心特色:
多维度标注:涵盖了标题、受众、标签、简介及集数。
AI 增强受众标签:针对部分原始数据未标注受众标签的问题,使用了专门的 sex_divide.py ,利用进行“TF-IDF + 逻辑回归”预测,并保留了预测置信度。
2. 数据字段说明 (Data Fields)
字段名
类型
说明
drama_id
string
脱敏后的剧集唯一编号 (例如 drama_0001)
title
string
短剧标题… See the full description on the dataset page: https://huggingface.co/datasets/EugeneMeng/Hongguo-Short-Drama-Corpus-AI-Labeled.cantonese-drama-voice-fine-grained语料集名称:面向影视剧AI配音的粤语语料库
语料来源:AI DimSum Lab
简介:
本语料库是专为粤语影视剧 AI 配音模型训练构建的专用语料资源,核心涵盖《神雕侠侣(1995 古天乐版)》《乘龙怪婿》《寻秦记》等经典粤语影视内容。语料库匹配 AI 配音模型的人物识别、语音情绪识别、语音生成三大模块需求,并根据下游任务对《神雕侠侣》等语音数据进行了多情感、多人物、文本标注。其数据规模总计约800MB,影视时长超 7 小时,可提供丰富的粤语语音、情感、人物关联样本,能有效支撑模型训练中人物区分、情感还原、语音生成的精度提升,是粤语影视剧 AI 配音落地的核心数据基础。
适用场景:
粤语影视剧 AI 配音模型训练:直接用于模型的人物区分、情感还原、语音生成模块优化;
粤语语音研究:可作为粤语语音特征、情感语音分析的基础数据集;
影视 AI 技术开发:为影视领域的语音合成、角色语音克隆等技术提供数据支持。
使用说明:
本语料库仅用于非商业研究与技术开发(商业使用需联系维护者确认授权);
使用前建议对语音数据进行预处理(如降噪、采样率统一),以提升模型训练效果。
ai-drama-production-harness-landscape-demo-media
Tin nhắn chưa gửi — landscape demo
Public media for the bundled 16:9 demo in AI Drama Production Harness.
Two fictional adult Vietnamese sisters, Mai and Linh, in two continuous dining-room scenes. AI-generated character/outfit/location references and synthetic voice references; 12 rendered clips plus the assembled film. No real-person reference recording is included.
The film is 93.197673 seconds, H.264/AAC, 864×480 (the workflow's rounded 480p preset). The seven reference… See the full description on the dataset page: https://huggingface.co/datasets/tungmtp/ai-drama-production-harness-landscape-demo-media.M-Drama
MDrama: Drama Video Understanding Dataset
Dataset Summary
32,361 QA annotations over 8,102 short-drama video clips (~52 GB, videos NOT included in this repo).
Each clip has ~4 annotations: 1 caption/summary + 3 QA (multiple-choice or open-ended).
Source videos are YouTube short dramas; use the url, start_time, end_time fields to retrieve each clip yourself.
Fields
Field
Description
qid
Unique question id (train_0 .. train_32360).… See the full description on the dataset page: https://huggingface.co/datasets/yixin1121/M-Drama.DramaticLibriQuoteChinese_drama_audioThis dataset is designed for the Emotional Speaking Style Retrieval (ESSR) task.
The file prompt.csv provides a detailed record of the natural language emotional description corresponding to each audio clip.
For the detail, please refer to: https://github.com/DeadWater1/FS-CLAP
jesus_dramasJesus Dramas is a collection of religious audio dramas across 430 languages. In total, there is around 640 hours of audio.
It can be used for language identification, spoken language modelling, or speech representation learning.
This dataset includes the raw unsegmented audio in a 16kHz single channel format. Each audio drama can have multiple speakers, for both male and female voices.
It can be segmented into utterances with a voice activity detection (VAD) model such as this one.
The… See the full description on the dataset page: https://huggingface.co/datasets/espnet/jesus_dramas.ears-dramabox
EARS — DramaBox Training Data
Multi-speaker DramaBox training dataset with 16,377 samples from the EARS (Expressive Anechoic Recordings of Speech) corpus. Features ~100+ diverse speakers with LLM-generated prompts.
Quick Facts
Property
Value
Samples
16,377
Speakers
~100+
Language
English
Format
WebDataset (.tar shards)
Shards
33 (500 samples each)
Source
EARS corpus
Prompt Generation
Gemma-3-4B-IT (creative, nuanced prompts)
Text… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/ears-dramabox.DramaTell_QA_v2_for_uploaddramabox-tuning-data
DramaBox Tuning Data
Paired training dataset with 12,330 samples designed for fine-tuning DramaBox with bidirectional audio pairs. Combines emotional speech (Emolia) and podcast data in a compact two-part format.
Quick Facts
Property
Value
Samples
12,330 total
Emolia subset
2,316 samples (emotional speech pairs)
Podcast subset
10,014 samples (diverse speaker pairs)
Languages
English, German, Spanish, French
Format
WebDataset (.tar shards)… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/dramabox-tuning-data.echo-robot-dramabox
Echo-TTS Robot — DramaBox-annotated (WER=0)
Robot-voice speech generated with Echo-TTS (EchoDiT preview, jordand/echo-tts-base)
using the Independent sampler preset (cfg_mode=independent, CFG=2,
speaker KV-scale=2), conditioned on a robot reference voice. Across the 40
EmoNet emotion categories, 5 utterances/emotion (200) were written by Gemini,
each generated with 5 seeds (1000 candidates), transcribed with
Parakeet-TDT-0.6B-v3, silence-aware trimmed (500 ms grace, cut in… See the full description on the dataset page: https://huggingface.co/datasets/ChristophSchuhmann/echo-robot-dramabox.DramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v2.0 (500 samples) - This release contains 500 carefully selected drama scripts with context-continuation pairs, designed to assess models across six independent evaluation dimensions. This represents a 5x expansion from v1.0, providing more comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/sxjiemust/DramaBench.podcast-dramabox-dacvae-pairs
podcast-dramabox-dacvae-pairs
Paired audio-codec latents for training a DramaBox → DACVAE latent translator.
Both codecs share an identical grid: 25 Hz, 128-dim, frame-aligned (same length).
Derived from TTS-AGI/podcast-tokenized-bg3.5-enj5.
How it was built (per sample)
DACVAE latent (from source dataset, = target) → DACVAE.decode → 48 kHz mono wav
→ duplicate to stereo → DramaBox/LTX-2.3 audio VAE encode → patchify → DramaBox latent (= input).
Both latents… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/podcast-dramabox-dacvae-pairs.imsdb-drama-screenplaydramallama-novels
DramaLlama dataset
This is the dataset repository of DramaLlama. This repository contains scripts designed to gather and prepare the dataset.
Note: This repository builds upon the findings of https://github.com/molbal/llm-text-completion-finetune
Step 1: Getting novels
We will use The Gutenberg project again to gather novels. Let's get some drama categories. I will aim for a larger dataset size this time.
I'm running the following scripts:
pip install requests… See the full description on the dataset page: https://huggingface.co/datasets/molbal/dramallama-novels.imsdb-drama-movie-scripts
Dataset Card for "imsdb-drama-movie-scripts"
More Information needed
DramaCV
Dataset Card for DramaCV
Dataset Summary
The DramaCV Dataset is an English-language dataset containing utterances of fictional characters in drama plays collected from Project Gutenberg. The dataset was automatically created by parsing 499 drama plays from the 15th to 20th century on Project Gutenberg, that are then parsed to attribute each character line to its speaker.
Task
This dataset was developed for Authorship Verification of literary characters. Each… See the full description on the dataset page: https://huggingface.co/datasets/gasmichel/DramaCV.big-drama-b23ef6
big-drama-b23ef6
Synthetic products test data: 32 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Elizabeth-Miller/big-drama-b23ef6.DRAMADRAMA-X
Request access to the original DRAMA dataset athttps://usa.honda-ri.com/drama#Downloadthedataset
Download the ZIP and extract the fileintegrated_output_v2.json
Ensure you have your existingdrama_x_annotations.json (with empty image_path/video_path fields) in the same folder.
Run the population script:
python populate_drama_x.py \
drama_x_annotated.json \
integrated_output_v2.json \
-o drama_x_annotations_populated.json
dramabox-gemini-finetune
DramaBox Gemini Fine-Tune Data
Pre-computed DramaBox latent-space training data with speaker embeddings for
fine-tuning with AdaLN-Zero speaker conditioning. All samples are same-speaker
part pairs (reference + target) with Gemini 3.5 Flash prompts.
Stats: 9,397 samples, 94 shards, 5.3 GB
Sample Structure
Each sample in the WebDataset tars contains:
File
Description
{key}_tgt.mp3
Target audio (MP3)
{key}_ref.mp3
Reference speaker audio (MP3)… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/dramabox-gemini-finetune.drama-booksHongguo-Short-Drama-Corpus-AI-Labeled
🎬 2025年红果短剧全量语料库 (AI标注版 V1)
Hongguo Short Drama Corpus with AI-Augmented Audience Labels
1. 数据集简介 (Dataset Summary)
本数据集包含了约 1500 条来自红果 (Hongguo/RedFruit) 平台的微短剧精选数据。本数据集旨在为中文短剧的 NLP 研究、市场趋势分析以及自动化剧名生成等任务提供高质量的基准数据。
核心特色:
多维度标注:涵盖了标题、受众、标签、简介及集数。
AI 增强受众标签:针对部分原始数据未标注受众标签的问题,使用了专门的 sex_divide.py ,利用进行“TF-IDF + 逻辑回归”预测,并保留了预测置信度。
2. 数据字段说明 (Data Fields)
字段名
类型
说明
drama_id
string
脱敏后的剧集唯一编号 (例如 drama_0001)
title
string… See the full description on the dataset page: https://huggingface.co/datasets/Oxiane/Hongguo-Short-Drama-Corpus-AI-Labeled.dramatic-box-5aeec9
dramatic-box-5aeec9
Synthetic products test data: 55 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/yeonghwanyang/dramatic-box-5aeec9.DramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v1.0 (100 samples) - This is the initial release containing 100 carefully selected drama scripts with context-continuation pairs, designed to assess models across six independent evaluation dimensions.
Release Roadmap
Version
Samples
Status… See the full description on the dataset page: https://huggingface.co/datasets/LizRob6913/DramaBench.dramallama-novels
DramaLlama dataset
This is the dataset repository of DramaLlama. This repository contains scripts designed to gather and prepare the dataset.
Note: This repository builds upon the findings of https://github.com/molbal/llm-text-completion-finetune
Step 1: Getting novels
We will use The Gutenberg project again to gather novels. Let's get some drama categories. I will aim for a larger dataset size this time.
I'm running the following scripts:
pip install requests… See the full description on the dataset page: https://huggingface.co/datasets/LizRob6913/dramallama-novels.DramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v1.0 (100 samples) - This is the initial release containing 100 carefully selected drama scripts with context-continuation pairs, designed to assess models across six independent evaluation dimensions.
Release Roadmap
Version
Samples
Status… See the full description on the dataset page: https://huggingface.co/datasets/tatan2/DramaBench.
