datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu-zh-chengyu-cpt
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus
A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on
cultural knowledge in figurative language, built from the highest-quality
tier of opencsg/Fineweb-Edu-Chinese-V2.1.
Each document is educational Chinese text containing at least one culturally
vetted chengyu, with an appended 【成语注释】 knowledge block listing every
matched idiom's figurative meaning(s) and classical source citation.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.MiniShiftHER-Dataset
📚 HER-Dataset
Reasoning-Augmented Role-Playing Dataset for LLM Training
HER introduces dual-layer thinking that distinguishes characters' first-person thinking from LLMs' third-person thinking for cognitive-level persona simulation.
Overview
HER-Dataset is a high-quality role-playing dataset featuring reasoning-augmented dialogues extracted from literary works. The dataset includes:
📖 Rich character interactions from classic literature
🧠… See the full description on the dataset page: https://huggingface.co/datasets/ChengyuDu0123/HER-Dataset.chinese_traditional_chengyuGrasp_both
Bimanual Storage Task — Aligned Ego + UMI Dataset
Bimanual tabletop manipulation dataset with synchronized ego (head-mounted) and UMI (wrist-mounted) cameras. 50 episodes of a two-handed storage/organization task, each ~32 seconds.
Task
双手收纳 (Bimanual Storage): An operator uses two FastUMI Pro grippers to pick, move, and place objects on a tabletop. A head-mounted ego camera records a continuous third-person overhead view of the entire workspace.
Scene… See the full description on the dataset page: https://huggingface.co/datasets/chengyuanshu98/Grasp_both.chengyu_chinese
这是用于大模型微调的一个数据集,来源于COIG-CQIA里面的成语数据集。
##
仅用于学习使用。
llmail-inject-challenge
Dataset Summary
This dataset contains a large number of attack prompts collected as part of the now closed LLMail-Inject: Adaptive Prompt Injection Challenge.
We first describe the details of the challenge, and then we provide a documentation of the dataset
For the accompanying code, check out: https://github.com/microsoft/llmail-inject-challenge.
Citation
@article{abdelnabi2025,
title = {LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection… See the full description on the dataset page: https://huggingface.co/datasets/Chengyu22321/llmail-inject-challenge.
