datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chengyu時間:
2018年做成網站 https://chengyu.18dao.net
2024年用AI將文本生成圖片
2025年上傳到Hugging Face的Datasets
数据集中的文件总数: 20609
目录 "Text-to-Image/" 下的文件数量: 10296,子目錄數:5148,每個子目錄兩個文件,一個原始的文生圖png圖片,一個圖片解釋txt文件
目录 "image-chengyu/" 下的文件数量: 5155,加字的圖片jpg文件
目录 "text-chengyu/" 下的文件数量: 5156,文字解釋txt文件
M2AD
Visual Anomaly Detection under Complex View-Illumination Interplay: A Large-Scale Benchmark
🌐 Hugging Face Dataset
📚 Paper • 🏠 Homepageby Yunkang Cao*, Yuqi Cheng*, Xiaohao Xu, Yiheng Zhang, Yihan Sun, Yuxiang Tan, Yuxin Zhang, Weiming Shen,
🚀 Updates
We're committed to open science! Here's our progress:
2025/05/19: 📄 Paper released on ArXiv.
2025/05/16: 🌐 Dataset homepage launched.
2025/05/24: 🧪 Code release for benchmark evaluation! code… See the full description on the dataset page: https://huggingface.co/datasets/ChengYuQi99/M2AD.fineweb-edu-zh-chengyu-cpt
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus
A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on
cultural knowledge in figurative language, built from the highest-quality
tier of opencsg/Fineweb-Edu-Chinese-V2.1.
Each document is educational Chinese text containing at least one culturally
vetted chengyu, with an appended 【成语注释】 knowledge block listing every
matched idiom's figurative meaning(s) and classical source citation.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.MiniShiftHER-Dataset
📚 HER-Dataset
Reasoning-Augmented Role-Playing Dataset for LLM Training
HER introduces dual-layer thinking that distinguishes characters' first-person thinking from LLMs' third-person thinking for cognitive-level persona simulation.
Overview
HER-Dataset is a high-quality role-playing dataset featuring reasoning-augmented dialogues extracted from literary works. The dataset includes:
📖 Rich character interactions from classic literature
🧠… See the full description on the dataset page: https://huggingface.co/datasets/ChengyuDu0123/HER-Dataset.umi_raw_chase
UMI Raw Chase
这是 Chase 操作任务的原始 UMI 双手示范数据集。数据以 HDF5 文件保存,每个
episode_XX.hdf5 对应一条成功示范轨迹。
数据集概览
共 78 个 episodes、40,064 帧,大小约 850 MB。
采集频率为 30 Hz。
每条轨迹包含左右手腕 RGB 图像、左右手 FastUMI 位姿、左右夹爪输入和时间戳。
RGB 图像以 JPEG 压缩后存储,记录的图像尺寸属性为 320。
所有已收录 episode 的 success 属性均为 true。
目录与命名
目录
Episodes
帧数
说明
local_chase/
10
4,597
基础 Chase 数据
local_chase_disturb/
10
4,536
画面中存在不同颜色、形状的干扰物体
local_chase_fall/
5
3,170
目标物体抓取后跌落,再次抓取
local_chase_fallre/
3
2,139… See the full description on the dataset page: https://huggingface.co/datasets/chengyuanshu98/umi_raw_chase.chinese_traditional_chengyuGrasp_both
Bimanual Storage Task — Aligned Ego + UMI Dataset
Bimanual tabletop manipulation dataset with synchronized ego (head-mounted) and UMI (wrist-mounted) cameras. 50 episodes of a two-handed storage/organization task, each ~32 seconds.
Task
双手收纳 (Bimanual Storage): An operator uses two FastUMI Pro grippers to pick, move, and place objects on a tabletop. A head-mounted ego camera records a continuous third-person overhead view of the entire workspace.
Scene… See the full description on the dataset page: https://huggingface.co/datasets/chengyuanshu98/Grasp_both.yidun_chengyuM2AD-datasetCIHPinstance-level_human_parsing
Images: images
Category_ids: semantic part segmentation labels Categories: visualized semantic part segmentation labels
Human_ids: semantic person segmentation labels Human: visualized semantic person segmentation labels
Instance_ids: instance-level human parsing labels Instances: visualized instance-level human parsing labels
Label order of semantic part segmentation:
1.Hat
2.Hair
3.Glove
4.Sunglasses
5.UpperClothes
6.Dress… See the full description on the dataset page: https://huggingface.co/datasets/chengyunlucky/CIHP.yidun_chengyuchengyu_chinese
这是用于大模型微调的一个数据集,来源于COIG-CQIA里面的成语数据集。
##
仅用于学习使用。
fastumi_raw_dataset_bottle
FastUMI Bottle Pick-and-Place Dataset (Raw)
Overview
45 episodes of bottle pick-and-place teleoperation data collected with FastUMI Pro on R1Pro (26') humanoid robot.
Task: Pick up a bottle from the table, place it at another location, return to start.
Six categories covering local/mobile/back modes and normal/fast speeds:
Category
Mode
Speed
Episodes
Total Frames
Total Duration
Avg Duration
local_standard
Local (base fixed)
Normal
10
4473
149.1s
14.9s… See the full description on the dataset page: https://huggingface.co/datasets/chengyuanshu98/fastumi_raw_dataset_bottle.llmail-inject-challenge
Dataset Summary
This dataset contains a large number of attack prompts collected as part of the now closed LLMail-Inject: Adaptive Prompt Injection Challenge.
We first describe the details of the challenge, and then we provide a documentation of the dataset
For the accompanying code, check out: https://github.com/microsoft/llmail-inject-challenge.
Citation
@article{abdelnabi2025,
title = {LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection… See the full description on the dataset page: https://huggingface.co/datasets/Chengyu22321/llmail-inject-challenge.
