aweffr/chinese-novel-continuation-precise-tokens
中文小说续写精确 Token 长度数据集 基于金庸《神雕侠侣》的高精度中文文本生成数据集,专为 GPU 内存测试和序列长度性能分析而设计。 数据集特点 高精度:99.4%+ 的目标 token 长度准确率 多种长度:5 个变体(1024、2048、4096、8192、16384 tokens) 统一格式:Alpaca 格式的小说续写任务 质量控制:95%+ 样本在目标长度的 ±2% 范围内 数据集统计 目标 Tokens 实际平均 准确率 样本数 文件大小 ±1% 内样本 ±2% 内样本 1024 1017.9 99.4% 800 2.5MB 739 775 2048 2037.7 99.5% 800 4.9MB 764 795 4096 4078.1 99.6% 800 9.8MB 792 800 8192 8158.1 99.6% 800 19.5MB 800 800 16384 16317.6 99.6% 800 39.1MB 800 800… See the full description on the dataset page: https://huggingface.co/datasets/aweffr/chinese-novel-continuation-precise-tokens.
This repository belongs to aweffr on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
