datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.deplyze-mini-dataset
Deplyze-Mini Dependency Intelligence Benchmark Dataset
This dataset contains standardized, ground-truth scenarios for training and evaluating software dependency intelligence models. It is designed to evaluate and prevent vulnerability hallucinations, train/test leakage, and prompt injection vulnerabilities in automated software composition analysis (SCA).
Dataset Composition
train.json: 400 multi-category dependency scenarios with instruction-tuning message… See the full description on the dataset page: https://huggingface.co/datasets/mosetireagan/deplyze-mini-dataset.MOSAIC
MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models
MOSAIC is a benchmark for evaluating the Moral, Social, and Individual dimensions of Large Language
Models across nine validated psychological questionnaires and four ethical-dilemma scenario sets.
This dataset accompanies the paper "MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large
Language Models" and the code at EricaCoppolillo/MOSAIC.
Dataset structure… See the full description on the dataset page: https://huggingface.co/datasets/EriCop/MOSAIC.MoS-Qwen3-8B-EAGLE3-responses
MoS — Qwen3-8B EAGLE3 Training Responses
Target-model responses for training EAGLE3 speculative-decoding draft models against
Qwen/Qwen3-8B. Built for the MoS (Mixture of
Speculators) project — a routed multi-MLP draft — and equally usable for any single-draft
EAGLE3 / SpecForge training run on Qwen3-8B.
599,087 complete assistant responses (with thinking traces) over five domains, generated
by Qwen3-8B itself so the draft learns to mimic the target's own distribution.… See the full description on the dataset page: https://huggingface.co/datasets/ryan-0608/MoS-Qwen3-8B-EAGLE3-responses.Slim-Moss003sft-zh因为原生的Moss003数量太大,所以进行了简单的去重。
去重方法大致为,只选择中文的对话,使用bert-base-chinese将第一个问题转换为embedding,使用类knn的方法抽取了1万条。并转换成了sharegpt格式。
TwinnyAI-Personas-Dataset
Overview
The TWINNY.AI Personas Dataset is a synthetic collection of 400 richly structured professional personas, engineered to power behavioral AI twins, persona-driven language model fine-tuning, and professional simulation systems.
Each persona is built from 14 attributes spanning demographics, professional context, behavioral psychology, and communication style sampled with realistic non-uniform distributions that mirror actual workforce demographics rather than uniform… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa190/TwinnyAI-Personas-Dataset.
