datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.mosaic-bench
MOSAIC
199 compositional attack chains across 10 real-world web applications, used to
benchmark whether AI coding agents will compose individually-routine tickets
into a deployable vulnerability.
Code & harness: https://github.com/mosaic-benchmark/mosaic-benchmark
Datasheet: DATASHEET.md · Croissant 1.1: croissant.json
What's in this release
Artifact
Contents
mosaic-bench.xlsx
Per-chain ASR (9 models, standard + resumed), BugBot verdicts (diff-mode +… See the full description on the dataset page: https://huggingface.co/datasets/MosaicBenchmark/mosaic-bench.19th-century-novelists19th-century novelists' sentences
We constructed the 5-author dataset using texts from Project Gutenberg, focusing on five prominent 19th-century novelists: Charles Dickens, Mark Twain, Herman Melville, Jane Austen, and Louisa May Alcott. This selection balances male and female authors as well as British and American literary traditions, offering a diverse testbed for stylistic analysis. Sentence segmentation was performed with the NLTK library, and tokenization/word counts were… See the full description on the dataset page: https://huggingface.co/datasets/Mosab-Rezaei/19th-century-novelists.chess-elite-uci
chess-elite-uci
A transformer-ready dataset of ~7.8 million elite chess games, pre-tokenized in UCI notation with a deterministic 1977-token vocabulary. Built for training chess language models directly with no preprocessing required.
Dataset Summary
Field
Value
Total games
7,805,503
Average sequence length
94.24 tokens
Max sequence length
255 tokens
Vocabulary size
1,977 tokens
Mean combined Elo
5,211 (~2,606 per player)
Sources… See the full description on the dataset page: https://huggingface.co/datasets/MostLime/chess-elite-uci.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json… See the full description on the dataset page: https://huggingface.co/datasets/lmdmengdi/moss-002-sft-data.mosaic
MOSAIC Dataset
This repository packages the public MOSAIC data artifacts from the paper "MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment."
MOSAIC is short for Multi-Objective Slice-Aware Iterative Curation for Alignment.
It contains three annotated source training pools and five training subsets selected by the MOSAIC search loop under a fixed 1M-token budget. The release also includes flattened iteration metadata so the search trajectory can be inspected… See the full description on the dataset page: https://huggingface.co/datasets/douyipu-real/mosaic.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/fuzhou-jiang/moss-002-sft-data.
