moe-expert
moe-7b-1b-active-shared-expertsgpt-oss-6.0b-specialized-all-pruned-moe-only-7-experts-i1-GGUFgpt-oss-6.0b-specialized-all-pruned-moe-only-7-experts-GGUFQwen1.5-MoE-A2.7B-20-experts-Maths-reasoning-trained-GGUFmoe-expert-coder-4bDeepseek-V2-13B-Math7K-Expert-Enhance-Subset-Expert-MoE-32-experts-i1-GGUFDeepseek-V2-13B-Math7K-Expert-Enhance-Subset-Expert-MoE-32-experts-GGUFQwen1.5-MoE-A2.7B-20-experts-SFT-trained-GGUF
gpt-oss-20b-moe-expert-power-traces-320k
GPT-OSS-20B MoE Expert Power Traces (320k, ChipWhisperer)
This dataset contains analog power traces captured with a ChipWhisperer Husky while running forced single-expert MoE computations derived from openai/gpt-oss-20b on an NVIDIA H100.
What is recorded
Each trace corresponds to one capture trial where:
A fixed expert id is selected (expert_00 ... expert_31).
A random hidden-state tensor is generated once per trial.
The selected expert computation is executed… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k.MoE_expert_selection_trace
📖 Introduction
This repository serves as a supplement to our paper "Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference".
It contains expert selection profiling traces of four top-tier MoE LLMs ranging from 235B to 1T (DeepSeek-R1, Kimi-K2-Thinking, Llama4-Marverick, and Qwen3-235B) across multiple benchmarks. For each query or request, we log the activated expert ID of every model layer of every generated token.
We provide analyses and… See the full description on the dataset page: https://huggingface.co/datasets/core12345/MoE_expert_selection_trace.gpt-oss-20b-moe-expert-power-traces-320k-ds16k
GPT-OSS-20B MoE Expert Power Traces (Downsampled to 16k)
Downsampled variant of the 320k expert-trace capture set.
Source
Raw source dataset (same captures):
32 experts (expert_00..expert_31)
10,000 traces per expert
320,000 total traces
raw trace length ~195k samples per trace
Downsampling method
Each raw trace was resampled to exactly 16384 samples using linear interpolation (np.interp) matching the trainer resampling step.
No baseline normalization and no… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k-ds16k.aime2024_Qwen3-30B-A3B_moe_patternsmath-500_Qwen3-30B-A3B_moe_patternsgpqa_diamond_Qwen3-30B-A3B_moe_patterns
