datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MoeGirlPedia_zh_cleaned_latest
🌐Language 中文|English
本数据集由2025年10月萌娘百科的快照经过清洗得来,专用于预训练等文本生成相关的模型训练。
特色
⚡体积优势
🧠文本易理解
💬更符合中文语境
仅经过基础清洗的数据集
1.06GB
存在复杂的网址链接残留的html标记正文内容被清除后残存的标题牛皮癣一样的引文注脚
暴力抹除非中文文字,导致信息缺失严重
本数据集
0.74GB(30.2%↓)
通过多重工序清洗基本不存在难以理解的文本内容保留部分英文以及少量其他语言文字(如日语)
仅经过基础清洗的数据集
size=66px|color=#8230FF|她已经不是我所认识的那个-{zh-hans:茜;zh-hant:仓式茜}-了。
'''仓式 茜'''(Kurashiki Akane)是由Spike Chunsoft所创作的系列游戏'''《极限脱出》'''及其衍生作品的主要角色之一。{{ZETOP}}
url=akanejunpei.jpg|position=up
图片说明=999中的茜(2027,21岁)
|本名=仓式 茜(くらしき… See the full description on the dataset page: https://huggingface.co/datasets/YCWTG/MoeGirlPedia_zh_cleaned_latest.Muice-Dataset
Muice-Dataset
沐雪角色扮演训练集
🤖ModelScope|
🤗HuggingFace|
(Github)Muicebot
更新日志
2026.05.18: 因为作者的论文使用到了本训练集需要引用,故更新 DOI 引用
2026.02.05: 小型更新,此次更新过后不再有新的数据集产生。
2025.08.23: 完整开源所有训练集以作研究用途,大幅更新自述文件
2025.02.14: 更新测试集以便透明化测试流程
2025.01.29: 新年快乐!为了感谢大家对沐雪训练集的喜欢,我们重写了训练集并额外提供 500 条训练集给大家。你可以在 这里 查看训练集重写目的和具体内容。除此之外,我们用 Sharegpt 格式规范了训练集格式,现在应该不会那么容易报错了...我们期望大家合理使用我们的训练集并训练出更高质量的模型,祝各位生活愉快。
简介… See the full description on the dataset page: https://huggingface.co/datasets/Moemu/Muice-Dataset.less-is-moe-s1-calibration-128-seq8192
Less-is-MoE S1K calibration data — 128 samples, seq_length 8192
This is the fixed calibration artifact used to prune GPT-OSS-120B,
Qwen3.5-122B-A10B, and the Gemma-4-26B-A4B causal language tower. It uses the same 128 source rows as the full-length variant:
yentinglin/s1K-1.1-trl-format revision
58a01564d278477da20ead1bcf1cde8e31f36251, train, followed by
Dataset.shuffle(seed=1234) and the first 128 nonempty messages rows.
For pruning, concatenate messages[].content with one… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128-seq8192.less-is-moe-s1-calibration-128
Less-is-MoE S1K calibration data — 128 full-length samples
This repository contains the exact 128 S1K rows selected for Less-is-MoE
full-model pruning. The selection reproduces the released loader:
source: yentinglin/s1K-1.1-trl-format
revision: 58a01564d278477da20ead1bcf1cde8e31f36251
split: train
order: Dataset.shuffle(seed=1234)
samples: first 128 nonempty messages rows
sequence-length limit: none
truncation: disabled
padding: disabled
calibration.jsonl stores every… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128.less-is-moe-gpqa-diamond-evaluation
Less-is-MoE GPQA-Diamond evaluation set
This private dataset stores the 198-question GPQA-Diamond evaluation file used
by the MoE-Honing evaluation format.
Upstream source: Idavidrein/gpqa, config gpqa_diamond
Upstream revision: 633f5ee89ab8ad4522a9f850766b73f62147ffdd
Split: test
Rows: 198
SHA-256: d5b0d6dad6c1993a8cb17fd7aa635fbc5e8b684ae42b5b6e80d15467b399eec3
Fields: problem, solution, domain
The problem field contains the formatted four-choice prompt, solution stores
the… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-gpqa-diamond-evaluation.arabic-bench-dataset
Arabic Bench Dataset
A curated evaluation dataset for benchmarking AI models on Arabic language tasks.
Overview
50 test cases across 8 categories
Each case includes a prompt, gold-standard reference answer, and a deliberately imperfect AI response
Covers Modern Standard Arabic (MSA) and multiple Arabic dialects
Designed for evaluating: translation, summarization, Q&A, creative writing, grammar, dialect understanding, legal/formal, and medical/scientific tasks… See the full description on the dataset page: https://huggingface.co/datasets/Moealsarraj/arabic-bench-dataset.
