datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SemDetect
NUpbr Ball Detection & Description Dataset
Synthetic renders of soccer balls (from the NUpbr generator)
with bounding boxes and free-text visual descriptions, for training semantic ball
detectors that generalise to ball appearances not seen during training.
Dataset structure
train/metadata.csv # 1440 images, 72 ball instances
train/*.png
validation/metadata.csv # 320 images, 16 ball instances
validation/*.png
test/metadata.csv # 300 images, 15… See the full description on the dataset page: https://huggingface.co/datasets/Ysobel/SemDetect.xiangqi-dataset
Xiangqi (Chinese Chess) Gameplay Trajectories & Visualizations Dataset
This dataset contains 1,000 high-quality Chinese Chess (Xiangqi) matches extracted and processed from the open-source training pipeline of Pikafish (the leading neural-network-backed Xiangqi engine). The original source training trajectories are credited to the px0data dataset on Kaggle.
For each match, this dataset provides both structured, step-by-step action sequences (JSONL format) suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/ysong18/xiangqi-dataset.robot_jersey_dataset
NUpbr RoboCup Jersey-Colour Robot Detection Dataset
1,707 physically-based-rendered images of RoboCup Humanoid Soccer scenes,
generated with NUpbr, each
containing 1-4 robots and one ball, with per-instance bounding boxes and,
for each robot, a jersey colour.
Every frame is restricted to at most two distinct jersey colours across all
robots present, regardless of robot count. Each robot is independently
assigned one of the two colours sampled fresh for that frame from a wide… See the full description on the dataset page: https://huggingface.co/datasets/Ysobel/robot_jersey_dataset.zh_spec_eval
中文专项评测集
本评测集共包含 512 条样本,分为 4 个工作负载(workload),每个工作负载包含 128 条样本。数据文件位于当前目录。
数据概览
Workload
来源数据集
数据划分
采样方法
Prompt 长度中位数(token)
zh_ceval
ceval/ceval-exam
val
汇总全部 52 个学科的样本,使用随机种子 0 打乱后取前 128 条,以兼顾学科覆盖的均衡性
96
zh_gaokao_math
hails/agieval-gaokao-mathqa
test(351 条)
使用随机种子 0 打乱后取前 128 条
142
zh_simpleqa
OpenStellarTeam/Chinese-SimpleQA
train(3,000 条)
使用随机种子 0 打乱后取前 128 条
30
zh_alpaca_gpt4
llm-wizard/alpaca-gpt4-data-zh
train(48,818 条)
过滤掉 instruction 少于… See the full description on the dataset page: https://huggingface.co/datasets/ysober/zh_spec_eval.MarineMISR
MarineMISR
MarineMISR is a multi-image super-resolution (MISR) dataset pairing stacks of
Landsat 8/9 scenes (low-resolution, 30 m) with a single co-located Sentinel-2
scene (high-resolution, 10 m) over coastal and marine habitats. Each sample is
a 512x512 pixel patch (5.12 km x 5.12 km) sampled from one of three habitat
types — coral reef, seagrass, and mangrove — so the dataset can be
used to train and evaluate super-resolution models specifically over these
ecologically… See the full description on the dataset page: https://huggingface.co/datasets/Ysobel/MarineMISR.
