datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xiangqi-dataset
Xiangqi (Chinese Chess) Gameplay Trajectories & Visualizations Dataset
This dataset contains 1,000 high-quality Chinese Chess (Xiangqi) matches extracted and processed from the open-source training pipeline of Pikafish (the leading neural-network-backed Xiangqi engine). The original source training trajectories are credited to the px0data dataset on Kaggle.
For each match, this dataset provides both structured, step-by-step action sequences (JSONL format) suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/ysong18/xiangqi-dataset.zh_spec_eval
中文专项评测集
本评测集共包含 512 条样本,分为 4 个工作负载(workload),每个工作负载包含 128 条样本。数据文件位于当前目录。
数据概览
Workload
来源数据集
数据划分
采样方法
Prompt 长度中位数(token)
zh_ceval
ceval/ceval-exam
val
汇总全部 52 个学科的样本,使用随机种子 0 打乱后取前 128 条,以兼顾学科覆盖的均衡性
96
zh_gaokao_math
hails/agieval-gaokao-mathqa
test(351 条)
使用随机种子 0 打乱后取前 128 条
142
zh_simpleqa
OpenStellarTeam/Chinese-SimpleQA
train(3,000 条)
使用随机种子 0 打乱后取前 128 条
30
zh_alpaca_gpt4
llm-wizard/alpaca-gpt4-data-zh
train(48,818 条)
过滤掉 instruction 少于… See the full description on the dataset page: https://huggingface.co/datasets/ysober/zh_spec_eval.
