datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hf-9routergpt-oss-20b-moe-expert-power-traces-320k
GPT-OSS-20B MoE Expert Power Traces (320k, ChipWhisperer)
This dataset contains analog power traces captured with a ChipWhisperer Husky while running forced single-expert MoE computations derived from openai/gpt-oss-20b on an NVIDIA H100.
What is recorded
Each trace corresponds to one capture trial where:
A fixed expert id is selected (expert_00 ... expert_31).
A random hidden-state tensor is generated once per trial.
The selected expert computation is executed… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k.OmniFake
OmniFake
OmniFake is a large-scale, well-categorized synthetic image dataset introduced in Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples. It contains 1.17 million AI-generated images from 45 distinct generators, paired with 1.17 million real images, designed for research on AI-generated image (AIGI) detection and source attribution.
For usage instructions and experimental protocols, please refer to the OmniDFA GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/MoeNew/OmniFake.moe-demo-clean
RoboTwin MOE Demo Clean
Raw RoboTwin demonstration data copied from bos:/lab-test/moe-demo-clean/.
The dataset is organized by task directories such as
place_can_basket-demo_clean-200/. Each task directory contains episode
subdirectories with an HDF5 trajectory file and an instructions.json file.
grug-moe-mix-swarm
Grug-MoE Data-Mix Experiments
The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below.
Fisher-DSP swarm (default)
840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2).
Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress
mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.story_clozespatial-moe-resultsATG-MoE_TrainingSet
ATG-MoE Pressure-Reducing Valve Assembly Training Set
This repository provides the training set used in our paper: "ATG-MoE: Autoregressive trajectory generation with mixture-of-experts for assembly skill learning".
The dataset is specifically designed for the pressure-reducing valve assembly task, featuring multi-skill robotic learning capabilities.
[!IMPORTANT]
This release only contains the training set.Evaluation requires a Unity-based simulation environment. The… See the full description on the dataset page: https://huggingface.co/datasets/failo0711/ATG-MoE_TrainingSet.FLAME-MoE-Traces
FLAME-MoE Routing Traces
Routing traces captured during pretraining of FLAME-MoE Mixture-of-Experts language models. For each token processed by the model, these traces record which experts the router selected (top-k expert IDs) and the corresponding gating probabilities (router softmax scores).
Architecture
Model
Params (Active/Total)
Transformer Layers
MoE Layers
Routed Experts
Shared Experts
Top-k
FLAME-MoE-290M
290M / 1.3B
9
8 (layers 2-9)
64
26
FLAME-MoE-721M
721M… See the full description on the dataset page: https://huggingface.co/datasets/CMU-FLAME/FLAME-MoE-Traces.moe-5g-nrxMoE_expert_selection_trace
📖 Introduction
This repository serves as a supplement to our paper "Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference".
It contains expert selection profiling traces of four top-tier MoE LLMs ranging from 235B to 1T (DeepSeek-R1, Kimi-K2-Thinking, Llama4-Marverick, and Qwen3-235B) across multiple benchmarks. For each query or request, we log the activated expert ID of every model layer of every generated token.
We provide analyses and… See the full description on the dataset page: https://huggingface.co/datasets/core12345/MoE_expert_selection_trace.moe-routing-drift-results
MoE routing drift — results
Measurements for a 2x2 experiment: adaptation (none / GEPA / prompt-tuning /
prefix-tuning) crossed with router retraining (frozen gate / gate retrained), on
inclusionAI/Ling-mini-2.0 and Qwen/Qwen3-30B-A3B-Instruct-2507. The weights are in
moe-routing-drift-checkpoints.
Content warning. quality/*/*.responses.jsonl contain verbatim comments from
civil_comments together with model outputs; the task is toxicity labelling, so the text
includes insults… See the full description on the dataset page: https://huggingface.co/datasets/AverageMetaheuristicsEnjoyer/moe-routing-drift-results.charlie-moe-nemotroncc-d004TinyMixtral-4x248M-MoE-atlas
juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas
A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing.
If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.Keural-MoE-14B-stage1-DatasetVideoVista_Train
VideoVista-Train
[📖 Paper] [📊 Dataset ] [✨ Github]
🌟 Citation
@article{li2024videovista,
title={Videovista: A versatile benchmark for video understanding and reasoning},
author={Li, Yunxin and Chen, Xinyu and Hu, Baotian and Wang, Longyue and Shi, Haoyuan and Zhang, Min},
journal={arXiv preprint arXiv:2406.11303},
year={2024}
}
🌟 Overview
VideoVista-Train consists of 114,581 training samples derived from 3,838 video clips.
These samples cover… See the full description on the dataset page: https://huggingface.co/datasets/Uni-MoE/VideoVista_Train.Holmes_moe_sft_dataMoE_datasetr1-671b-moemoe-speech
Recomendation of OOPPEENN's 56697375616C4E6F76656C5F44617461736574
I recommend OOPPEENN/56697375616C4E6F76656C5F44617461736574 for Japanese voice corpus, which is:
Similar speech domain to this one (Japanese anime-style speech from Japanese Visual Novel), but
Huge amounts of audio compared to this dataset (600 hours for this, 10,000 hours for Galgame_Dataset!)
This dataset contains about 50 games, and Galgame_Dataset contains more than 500 games!
Contains true transcripts of each… See the full description on the dataset page: https://huggingface.co/datasets/litagin/moe-speech.MoE-LLaVA
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
If you like our project, please give us a star ⭐ on GitHub for latest update.
📰 News
[2024.01.30] The paper is released.
[2024.01.27] 🤗Hugging Face demo and all codes & datasets are available now! Welcome to watch 👀 this repository for the latest updates.
😮 Highlights
MoE-LLaVA shows excellent performance in multi-modal learning.
🔥 High performance, but with fewer… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/MoE-LLaVA.VideoVista
📃 Paper | ✨ Project | 🏆 Leaderboard |
🌟 Citation
@article{li2024videovista,
title={Videovista: A versatile benchmark for video understanding and reasoning},
author={Li, Yunxin and Chen, Xinyu and Hu, Baotian and Wang, Longyue and Shi, Haoyuan and Zhang, Min},
journal={arXiv preprint arXiv:2406.11303},
year={2024}
}
Overview
The JSON file contains all video QA pairs (about 25,000).
The merged.zip* files consist of all sourced videos (3402).
The… See the full description on the dataset page: https://huggingface.co/datasets/Uni-MoE/VideoVista.MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials.
Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset.
Looking forward to see more models and synthetic datasets trained from this raw archive, good luck!
Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.MSVQAUnofficial training-ready fork of Kaij00/MSVQA.
gpt-oss-20b-moe-expert-power-traces-320k-ds16k
GPT-OSS-20B MoE Expert Power Traces (Downsampled to 16k)
Downsampled variant of the 320k expert-trace capture set.
Source
Raw source dataset (same captures):
32 experts (expert_00..expert_31)
10,000 traces per expert
320,000 total traces
raw trace length ~195k samples per trace
Downsampling method
Each raw trace was resampled to exactly 16384 samples using linear interpolation (np.interp) matching the trainer resampling step.
No baseline normalization and no… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k-ds16k.Muice-Dataset
Muice-Dataset
沐雪角色扮演训练集
🤖ModelScope|
🤗HuggingFace|
(Github)Muicebot
更新日志
2026.05.18: 因为作者的论文使用到了本训练集需要引用,故更新 DOI 引用
2026.02.05: 小型更新,此次更新过后不再有新的数据集产生。
2025.08.23: 完整开源所有训练集以作研究用途,大幅更新自述文件
2025.02.14: 更新测试集以便透明化测试流程
2025.01.29: 新年快乐!为了感谢大家对沐雪训练集的喜欢,我们重写了训练集并额外提供 500 条训练集给大家。你可以在 这里 查看训练集重写目的和具体内容。除此之外,我们用 Sharegpt 格式规范了训练集格式,现在应该不会那么容易报错了...我们期望大家合理使用我们的训练集并训练出更高质量的模型,祝各位生活愉快。
简介… See the full description on the dataset page: https://huggingface.co/datasets/Moemu/Muice-Dataset.lewm-vjepa21-state-moe-shared-residual-wo-tokencl-analysis
LEWM V-JEPA2.1 State-MoE Shared-Residual Analysis
This dataset contains PNG visualizations and adjacent-epoch absolute-delta
summaries for the experiment
lewm_reasoning_vjepa21_vitL_tokens_100tasks_stateMoE_sharedResidual_pair_topkExcess_flat_headaware_woTokenCL_modify.
Contents
metadata.csv: parsed stage, epoch, layer, condition, scope, and image paths.
gallery_manifest.json: manifest consumed by the companion Static HTML Space.
summary.json: aggregate image… See the full description on the dataset page: https://huggingface.co/datasets/zoeloopy/lewm-vjepa21-state-moe-shared-residual-wo-tokencl-analysis.aime2024_Qwen3-30B-A3B_moe_patternsmath-500_Qwen3-30B-A3B_moe_patternsVideoVista-Event
🌟 Citation
If you find our data is useful for your work, please cite our work:
@article{li2024videovista,
title={Videovista: A versatile benchmark for video understanding and reasoning},
author={Li, Yunxin and Chen, Xinyu and Hu, Baotian and Wang, Longyue and Shi, Haoyuan and Zhang, Min},
journal={arXiv preprint arXiv:2406.11303},
year={2024}
}
This dataset provides sequences of video event descriptions, automatically generated by our annotation framework.
The code for… See the full description on the dataset page: https://huggingface.co/datasets/Uni-MoE/VideoVista-Event.
