datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chinese-small-lm-corpus
Chinese Small LM Corpus
用于中文小型语言模型预训练的统一文本语料,字段为:
text:规范化后的训练文本
source:原始数据集名称
数据量
来源
有效样本数
TinyStories-Zh-2M
1,994,291
Wikipedia-20231101.zh
1,384,748
Zhihu-KOL
1,002,863
总计
4,381,902
来源与许可
RobinChen2001/TinyStories-Zh-2M:数据卡标注 MIT;同时应检查英文上游数据及机器翻译来源条款。
wikimedia/wikipedia (20231101.zh):CC BY-SA 3.0 与 GFDL。
wangrui6/Zhihu-KOL:原数据卡未声明许可证。
此合并数据集不提供统一的再授权。下载者须分别遵守各来源的许可、署名、隐私与内容使用要求。
lmarena-100k-long-sample-prompts-completions-Mistral-Small-24B-Instruct-2501Small-LM-PretrainingVNTL-v2-2k-small
Dataset Card for "VNTL-v2-2k-small"
More Information needed
ZERO-lm-dataset-chat-smallA combination of multiple open datasets combined into one for training simple lms.
small_lmd_midi
