drama
Datasets
All datasets matching “drama”DramaADDramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v3.0 Full (1,103 samples) - The complete DramaBench collection is now openly available, with context-continuation pairs designed to assess models across six independent evaluation dimensions.
Release Roadmap
Version
Samples
Status… See the full description on the dataset page: https://huggingface.co/datasets/FutureMa/DramaBench.ai-drama-production-harness-demo-media
The Second Key — demo media
This local staging tree contains reused AI-generated reference art and
model-rendered video from the operator-owned project run.
QC disclosure
The operator explicitly skipped measured per-clip QC for this set
(projects/default/runs/qc_skipped.json, schema version 1,
reason operator_satisfied). This is an operator QC waiver, not a fabricated
passing QC verdict. The staging gate independently resolved the current render
source and… See the full description on the dataset page: https://huggingface.co/datasets/tungmtp/ai-drama-production-harness-demo-media.Chinese_drama_audioThis dataset is designed for the Emotional Speaking Style Retrieval (ESSR) task.
The file prompt.csv provides a detailed record of the natural language emotional description corresponding to each audio clip.
For the detail, please refer to: https://github.com/DeadWater1/FS-CLAP
Hongguo-Short-Drama-Corpus-AI-Labeled
🎬 2025年红果短剧全量语料库 (AI标注版 V1)
Hongguo Short Drama Corpus with AI-Augmented Audience Labels
1. 数据集简介 (Dataset Summary)
本数据集包含了约 1500 条来自红果 (Hongguo/RedFruit) 平台的微短剧精选数据。本数据集旨在为中文短剧的 NLP 研究、市场趋势分析以及自动化剧名生成等任务提供高质量的基准数据。
核心特色:
多维度标注:涵盖了标题、受众、标签、简介及集数。
AI 增强受众标签:针对部分原始数据未标注受众标签的问题,使用了专门的 sex_divide.py ,利用进行“TF-IDF + 逻辑回归”预测,并保留了预测置信度。
2. 数据字段说明 (Data Fields)
字段名
类型
说明
drama_id
string
脱敏后的剧集唯一编号 (例如 drama_0001)
title
string
短剧标题… See the full description on the dataset page: https://huggingface.co/datasets/EugeneMeng/Hongguo-Short-Drama-Corpus-AI-Labeled.cantonese-drama-voice-fine-grained语料集名称:面向影视剧AI配音的粤语语料库
语料来源:AI DimSum Lab
简介:
本语料库是专为粤语影视剧 AI 配音模型训练构建的专用语料资源,核心涵盖《神雕侠侣(1995 古天乐版)》《乘龙怪婿》《寻秦记》等经典粤语影视内容。语料库匹配 AI 配音模型的人物识别、语音情绪识别、语音生成三大模块需求,并根据下游任务对《神雕侠侣》等语音数据进行了多情感、多人物、文本标注。其数据规模总计约800MB,影视时长超 7 小时,可提供丰富的粤语语音、情感、人物关联样本,能有效支撑模型训练中人物区分、情感还原、语音生成的精度提升,是粤语影视剧 AI 配音落地的核心数据基础。
适用场景:
粤语影视剧 AI 配音模型训练:直接用于模型的人物区分、情感还原、语音生成模块优化;
粤语语音研究:可作为粤语语音特征、情感语音分析的基础数据集;
影视 AI 技术开发:为影视领域的语音合成、角色语音克隆等技术提供数据支持。
使用说明:
本语料库仅用于非商业研究与技术开发(商业使用需联系维护者确认授权);
使用前建议对语音数据进行预处理(如降噪、采样率统一),以提升模型训练效果。
