datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Muice-Dataset
Muice-Dataset
沐雪角色扮演训练集
🤖ModelScope|
🤗HuggingFace|
(Github)Muicebot
更新日志
2026.05.18: 因为作者的论文使用到了本训练集需要引用,故更新 DOI 引用
2026.02.05: 小型更新,此次更新过后不再有新的数据集产生。
2025.08.23: 完整开源所有训练集以作研究用途,大幅更新自述文件
2025.02.14: 更新测试集以便透明化测试流程
2025.01.29: 新年快乐!为了感谢大家对沐雪训练集的喜欢,我们重写了训练集并额外提供 500 条训练集给大家。你可以在 这里 查看训练集重写目的和具体内容。除此之外,我们用 Sharegpt 格式规范了训练集格式,现在应该不会那么容易报错了...我们期望大家合理使用我们的训练集并训练出更高质量的模型,祝各位生活愉快。
简介… See the full description on the dataset page: https://huggingface.co/datasets/Moemu/Muice-Dataset.moe-unified-dataset-sota
moe-unified-dataset-sota
A unified dataset for training Mixture of Experts (MoE) models, combining multiple high-quality sources.
Dataset Statistics
Total Examples: 2,186,763
Train Split: 2,077,424
Test Split: 109,339
Sources
NousResearch/Hermes-3-Dataset - General instruction following, math, coding (~950k examples)
Salesforce/xlam-function-calling-60k - Function/tool calling (60k examples)
MegaScience/TextbookReasoning - Academic Q&A (~650k examples)… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/moe-unified-dataset-sota.MOE-RMCD
ytchen175/MOE-RMCD
一個精心設計的繁體中文 / 正體中文指令資料集。
A delicate Traditional Chinese instructions following dataset.
Introduction
「教育部重編國語辭典修訂本指令資料集」 (Ministry of Education Revised Mandarin Chinese Dictionary Instruction Dataset,簡稱 MOE-RMCD),
它是由教育部的《重編國語辭典修訂本》為底所構建的指令資料集。
基於想要盡可能最大化利用原始資料潛在價值的想法,
我們從中抽取出五大類任務 ── 詞語解釋、簡繁轉換、單句釋義、近似詞與反義詞,共計 36 萬筆指令 (instructions)。
更詳細的資料處理流程請見 1.3_preprocess_教育部重編國語辭典修訂本.ipynb。
Scope
排除了過於罕見的字,只留下中日韓統一表意文字列表 (CJK Unified Ideographs) 與… See the full description on the dataset page: https://huggingface.co/datasets/ytchen175/MOE-RMCD.less-is-moe-gpqa-diamond-evaluation
Less-is-MoE GPQA-Diamond evaluation set
This private dataset stores the 198-question GPQA-Diamond evaluation file used
by the MoE-Honing evaluation format.
Upstream source: Idavidrein/gpqa, config gpqa_diamond
Upstream revision: 633f5ee89ab8ad4522a9f850766b73f62147ffdd
Split: test
Rows: 198
SHA-256: d5b0d6dad6c1993a8cb17fd7aa635fbc5e8b684ae42b5b6e80d15467b399eec3
Fields: problem, solution, domain
The problem field contains the formatted four-choice prompt, solution stores
the… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-gpqa-diamond-evaluation.MoeGirlQA
数据集名称
MoeGirlQA
数据集描述
本数据集是一个通过大语言模型(LLM)批量生成的问答对(QA)集合。其原始文本内容来源于萌娘百科。
数据生成方法
原始内容:从萌娘百科的原始页面中提取文本信息。
自动生成:将原始文本输入给大语言模型,指令其根据内容批量生成相关的问答对。
处理说明:生成过程为全自动化,未经过人工校对与修正。
已知问题与使用限制
请使用者特别注意:
可能存在事实性错误(幻觉):由于问答对由LLM自动生成且未经人工校验,答案中可能包含与原始百科内容不符、编造或扭曲的信息。本数据集不保证其事实准确性。
覆盖范围不完整:本数据集不承诺覆盖萌娘百科的所有条目或某个条目的全部内容,生成范围受提取的原始文本和模型能力限制。
使用许可
本数据集基于 Creative Commons Attribution-NonCommercial-ShareAlike 3.0 (CC BY-NC-SA 3.0) 协议发布。
相关链接… See the full description on the dataset page: https://huggingface.co/datasets/cyberlangke/MoeGirlQA.arabic-bench-dataset
Arabic Bench Dataset
A curated evaluation dataset for benchmarking AI models on Arabic language tasks.
Overview
50 test cases across 8 categories
Each case includes a prompt, gold-standard reference answer, and a deliberately imperfect AI response
Covers Modern Standard Arabic (MSA) and multiple Arabic dialects
Designed for evaluating: translation, summarization, Q&A, creative writing, grammar, dialect understanding, legal/formal, and medical/scientific tasks… See the full description on the dataset page: https://huggingface.co/datasets/Moealsarraj/arabic-bench-dataset.
