datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime2024_Qwen3-30B-A3B_moe_patternsmath-500_Qwen3-30B-A3B_moe_patternsgpqa_diamond_Qwen3-30B-A3B_moe_patternsall_datasets_Qwen3-30B-A3B_moe_patternsuniagent-qwen3-30b-a3b-r2e-rollouts-r2e_moe_maxrl_09111923mmlu_Qwen3-30B-A3B_moe_patternswmt16_Qwen3-30B-A3B_moe_patternsMoE-Router-Dataset-MMMLU-Qwen3-30B-A3B
Dataset
Qwen3-30B-A3B 모델에 MMLU와 MMMLU의 영어/한국어 데이터를 넣고, gate가 선정한 top 8 expert의 id를 추출했습니다.
think/nonthink 모드 둘 다 생성했습니다.
생성 하이퍼파라미터
max_prompt_tokens = 2048 # MMMLU 최대 프롬프트 토큰: 1500+
max_think_tokens = 1024
max_nonthink_tokens = 1024
temperature = 0.6
top_p = 0.95
생성 소스코드: https://github.com/werty1248/MoE-Analyzer-vLLM
aime2024_Qwen3-30B-A3B_moe_patterns_logitsxsum_Qwen3-30B-A3B_moe_patternsMoE-Router-Dataset-Statistic-MMMLU-Qwen3-30B-A3B
werty1248/MoE-Router-Dataset-MMMLU-Qwen3-30B-A3B 데이터에서 expert 통계를 낸 데이터
prompt_stat: 프롬프트 처리 시 선택된 expert id 통계
output_stat: 토큰 생성 시 선택된 expert id 통계
first_stat: 생성된 토큰 중 최초 128토큰에서 선택된 expert id 통계
last_stat: 생성된 토큰 중 마지막 128토큰에서 선택된 expert id 통계
total_stat: prompt_stat + output_stat
Dataset
Qwen3-30B-A3B 모델에 MMLU와 MMMLU의 영어/한국어 데이터를 넣고, gate가 선정한 top 8 expert의 id를 추출했습니다.
think/nonthink 모드 둘 다 생성했습니다.
생성 하이퍼파라미터
max_prompt_tokens = 2048 # MMMLU 최대 프롬프트 토큰: 1500+… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/MoE-Router-Dataset-Statistic-MMMLU-Qwen3-30B-A3B.math_datasets_Qwen3-30B-A3B_moe_patternsmath-500_Qwen3-30B-A3B_moe_patterns_logitsqwen3_moe_30A3B_instr_2507_mt-benchdeepscaler_qwen3_4b_8r_all_wrong_qwen3p5_moe_397b_math_thinkingdapo_math_17k_qwen3_4b_8r_all_wrong_qwen3p5_moe_397b_math_thinking
