datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.dolly_hhrlhf
Dataset Card for "dolly_hhrlhf"
This dataset is a combination of Databrick's dolly-15k dataset and a filtered subset of Anthropic's HH-RLHF. It also includes a test split, which was missing in the original dolly set. That test set is composed of 200 randomly selected samples from dolly + 4,929 of the test set samples from HH-RLHF which made it through the filtering process. The train set contains 59,310 samples; 15,014 - 200 = 14,814 from Dolly, and the remaining 44,496 from… See the full description on the dataset page: https://huggingface.co/datasets/mosaicml/dolly_hhrlhf.moshub-code
Mos.Hub Code Dataset
Dataset Description
This dataset was compiled from code repositories hosted on Mos.Hub (hub.mos.ru), a code hosting platform operated by the Moscow Government. Mos.Hub is a service for storing and working with source code, based on the Git version control system, primarily used by Russian developers and government-related projects.
Dataset Summary
Statistic
Value
Total Files
15,740,580
Total Repositories
16,130… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/moshub-code.LiveClawbench-trajectoriesLiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
Overview
LLM agents are increasingly expected to handle real-world assistant tasks — booking flights, managing emails, debugging code, curating knowledge bases — yet existing benchmarks evaluate them under isolated difficulty sources. LiveClawBench addresses this gap by introducing a Triple-Axis Complexity Framework and building a benchmark of 134 manually constructed tasks with explicit factor… See the full description on the dataset page: https://huggingface.co/datasets/Mosi-AI/LiveClawbench-trajectories.mosaic-bench
MOSAIC
199 compositional attack chains across 10 real-world web applications, used to
benchmark whether AI coding agents will compose individually-routine tickets
into a deployable vulnerability.
Code & harness: https://github.com/mosaic-benchmark/mosaic-benchmark
Datasheet: DATASHEET.md · Croissant 1.1: croissant.json
What's in this release
Artifact
Contents
mosaic-bench.xlsx
Per-chain ASR (9 models, standard + resumed), BugBot verdicts (diff-mode +… See the full description on the dataset page: https://huggingface.co/datasets/MosaicBenchmark/mosaic-bench.19th-century-novelists19th-century novelists' sentences
We constructed the 5-author dataset using texts from Project Gutenberg, focusing on five prominent 19th-century novelists: Charles Dickens, Mark Twain, Herman Melville, Jane Austen, and Louisa May Alcott. This selection balances male and female authors as well as British and American literary traditions, offering a diverse testbed for stylistic analysis. Sentence segmentation was performed with the NLTK library, and tokenization/word counts were… See the full description on the dataset page: https://huggingface.co/datasets/Mosab-Rezaei/19th-century-novelists.MOSAIC
MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models
MOSAIC is a benchmark for evaluating the Moral, Social, and Individual dimensions of Large Language
Models across nine validated psychological questionnaires and four ethical-dilemma scenario sets.
This dataset accompanies the paper "MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large
Language Models" and the code at EricaCoppolillo/MOSAIC.
Dataset structure… See the full description on the dataset page: https://huggingface.co/datasets/EriCop/MOSAIC.mosaic-emnlp2026
MOSAIC
Dataset Summary
MOSAIC is a course-centric multimodal dataset accompanying a Findings of EMNLP 2026 paper. The dataset centers on mosaic.jsonl, a JSONL file that stores course-level metadata together with nested video-level summaries, subtitles, captions, and auxiliary references.
The public release also includes:
data/graph_p_results/: course-level knowledge graph JSON files keyed by kg
data/all.csv: source-resource metadata for video-level slide… See the full description on the dataset page: https://huggingface.co/datasets/AmyIvan/mosaic-emnlp2026.deplyze-mini-dataset
Deplyze-Mini Dependency Intelligence Benchmark Dataset
This dataset contains standardized, ground-truth scenarios for training and evaluating software dependency intelligence models. It is designed to evaluate and prevent vulnerability hallucinations, train/test leakage, and prompt injection vulnerabilities in automated software composition analysis (SCA).
Dataset Composition
train.json: 400 multi-category dependency scenarios with instruction-tuning message… See the full description on the dataset page: https://huggingface.co/datasets/mosetireagan/deplyze-mini-dataset.chess-elite-uci
chess-elite-uci
A transformer-ready dataset of ~7.8 million elite chess games, pre-tokenized in UCI notation with a deterministic 1977-token vocabulary. Built for training chess language models directly with no preprocessing required.
Dataset Summary
Field
Value
Total games
7,805,503
Average sequence length
94.24 tokens
Max sequence length
255 tokens
Vocabulary size
1,977 tokens
Mean combined Elo
5,211 (~2,606 per player)
Sources… See the full description on the dataset page: https://huggingface.co/datasets/MostLime/chess-elite-uci.MoS-Qwen3-8B-EAGLE3-responses
MoS — Qwen3-8B EAGLE3 Training Responses
Target-model responses for training EAGLE3 speculative-decoding draft models against
Qwen/Qwen3-8B. Built for the MoS (Mixture of
Speculators) project — a routed multi-MLP draft — and equally usable for any single-draft
EAGLE3 / SpecForge training run on Qwen3-8B.
599,087 complete assistant responses (with thinking traces) over five domains, generated
by Qwen3-8B itself so the draft learns to mimic the target's own distribution.… See the full description on the dataset page: https://huggingface.co/datasets/ryan-0608/MoS-Qwen3-8B-EAGLE3-responses.mOSCAR-Persian
Dataset Card for mOSCAR-Persian
This is a clone of mOSCAR Persian split (images excluded), which is further divided into single documents. Both URLs and Document IDs are consistent with mOSCAR.
Dataset Sources
Repository: mOSCAR
Paper [optional]: mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
Citation
@article{futeral2024moscar,
title={mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus},
author={Futeral… See the full description on the dataset page: https://huggingface.co/datasets/ErfanMoosaviMonazzah/mOSCAR-Persian.arabic-agent-eval
Arabic Agent Eval — Dataset Card
An open, installable Arabic function-calling benchmark with dialect splits.
Dataset summary
51 evaluation items spanning 6 categories and 5 dialects of Arabic, testing whether large language models can (a) select the right tool, (b) extract arguments from natural Arabic instructions, (c) preserve Arabic text in tool arguments instead of transliterating, and (d) understand dialectal framing.
Supported tasks… See the full description on the dataset page: https://huggingface.co/datasets/Mosescreates/arabic-agent-eval.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json… See the full description on the dataset page: https://huggingface.co/datasets/lmdmengdi/moss-002-sft-data.moshub
Mos.Hub Code Dataset
A comprehensive code dataset compiled from Mos.Hub, Moscow's official code hosting platform operated by the Moscow Government. This dataset is designed to support training code models with authentic Russian development practices and documentation.
Overview
The Mos.Hub Code Dataset represents a significant code corpus from Russia's governmental and municipal code hosting platform, capturing diverse projects across 297 programming languages. It… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/moshub.moscar_urls
Dataset Card for moscar_urls
This dataset provides the URLs and top-level domains associated with training records in oscar-corpus/mOSCAR. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/moscar_urls.mosaic
MOSAIC Dataset
This repository packages the public MOSAIC data artifacts from the paper "MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment."
MOSAIC is short for Multi-Objective Slice-Aware Iterative Curation for Alignment.
It contains three annotated source training pools and five training subsets selected by the MOSAIC search loop under a fixed 1M-token budget. The release also includes flattened iteration metadata so the search trajectory can be inspected… See the full description on the dataset page: https://huggingface.co/datasets/douyipu-real/mosaic.mimi_tokenizer
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mosh2i/mimi_tokenizer.fatima_blind_spot_challengeGot it. From now on I'll write everything inside Markdown blocks so you can copy easily.
Here is your full content entirely in Markdown:
# Fatima Fellowship 2026: Technical Challenge - Model Blind Spots
## 1. Model Overview
- **Model Tested:** [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Base Model)
- **Parameters:** 0.6B
- **Type:** Causal Language Model (Base / Pre-trained)
---
## 2. Methodology & Loading
To evaluate the model, I used **Google Colab** with a **T4… See the full description on the dataset page: https://huggingface.co/datasets/moseleydev/fatima_blind_spot_challenge.Slim-Moss003sft-zh因为原生的Moss003数量太大,所以进行了简单的去重。
去重方法大致为,只选择中文的对话,使用bert-base-chinese将第一个问题转换为embedding,使用类knn的方法抽取了1万条。并转换成了sharegpt格式。
moscar-corpus-thai-cleaned
Dataset mOSCAR Thai Cleaned
ชุดข้อมูลนี้เป็นชุดข้อมูลภาษาไทยขนาดใหญ่ที่ผ่านการทำความสะอาดแล้ว เหมาะสำหรับงานประมวลผลภาษาธรรมชาติ (NLP) เช่น การฝึกสอนโมเดลภาษา การสรุปผล การแปลภาษา ฯลฯ
รายละเอียดชุดข้อมูล
จำนวนตัวอย่าง: 1,643,471 ตัวอย่าง (train)
ขนาดข้อมูล: 5,132,779,656 ไบต์
ฟีเจอร์:
title (string): หัวข้อหรือข้อความแรกของแต่ละตัวอย่าง
text (string): เนื้อหาข้อความภาษาไทยที่ผ่านการคัดกรองและทำความสะอาดแล้ว
ภาษา: ไทย (th)
ลิขสิทธิ์: Apache-2.0
ขนาด: 1M < n < 10M… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/moscar-corpus-thai-cleaned.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/fuzhou-jiang/moss-002-sft-data.TwinnyAI-Personas-Dataset
Overview
The TWINNY.AI Personas Dataset is a synthetic collection of 400 richly structured professional personas, engineered to power behavioral AI twins, persona-driven language model fine-tuning, and professional simulation systems.
Each persona is built from 14 attributes spanning demographics, professional context, behavioral psychology, and communication style sampled with realistic non-uniform distributions that mirror actual workforce demographics rather than uniform… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa190/TwinnyAI-Personas-Dataset.Moshpit-Combined-R2-Uncensored
MoshPIT Is to be Archived By the end of November, Please More Updated Corpus Like Iris, This Dataset Offer No Benefit and Provides alot of Hallucination
📜 Please read our Terms and Conditions before using this dataset.
MoshPIT-R2 COMBINED UNCENSORED
MoshPIT R2, is a dataset combining multiple Moshpit Iteration to make a 500 Dataset examples,
MoshPIT R2 is purely generated through GPT2-XL,
Mushed-R1 Dataset is Synthetic, It has no Curation But can offer Quick Dataset… See the full description on the dataset page: https://huggingface.co/datasets/N-Bot-Int/Moshpit-Combined-R2-Uncensored.
