CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AverageMetaheuristicsEnjoyer /moe-routing-drift-results MoE routing drift — results Measurements for a 2x2 experiment: adaptation (none / GEPA / prompt-tuning / prefix-tuning) crossed with router retraining (frozen gate / gate retrained), on inclusionAI/Ling-mini-2.0 and Qwen/Qwen3-30B-A3B-Instruct-2507. The weights are in moe-routing-drift-checkpoints. Content warning. quality/*/*.responses.jsonl contain verbatim comments from civil_comments together with model outputs; the task is toxicity labelling, so the text includes insults… See the full description on the dataset page: https://huggingface.co/datasets/AverageMetaheuristicsEnjoyer/moe-routing-drift-results.tabulartext-classificationn<1K1 likes970 downloads11h agoHugging Face02YCWTG /MoeGirlPedia_zh_cleaned_latest 🌐Language 中文|English 本数据集由2025年10月萌娘百科的快照经过清洗得来,专用于预训练等文本生成相关的模型训练。 特色 ⚡体积优势 🧠文本易理解 💬更符合中文语境 仅经过基础清洗的数据集 1.06GB 存在复杂的网址链接残留的html标记正文内容被清除后残存的标题牛皮癣一样的引文注脚 暴力抹除非中文文字,导致信息缺失严重 本数据集 0.74GB(30.2%↓) 通过多重工序清洗基本不存在难以理解的文本内容保留部分英文以及少量其他语言文字(如日语) 仅经过基础清洗的数据集 size=66px|color=#8230FF|她已经不是我所认识的那个-{zh-hans:茜;zh-hant:仓式茜}-了。 '''仓式 茜'''(Kurashiki Akane)是由Spike Chunsoft所创作的系列游戏'''《极限脱出》'''及其衍生作品的主要角色之一。{{ZETOP}} url=akanejunpei.jpg|position=up 图片说明=999中的茜(2027,21岁) |本名=仓式 茜(くらしき… See the full description on the dataset page: https://huggingface.co/datasets/YCWTG/MoeGirlPedia_zh_cleaned_latest.texttext-generation100K<n<1M4 likes344 downloads11mo agoHugging Face03Moemu /Muice-Dataset Muice-Dataset 沐雪角色扮演训练集 🤖ModelScope| 🤗HuggingFace| (Github)Muicebot 更新日志 2026.05.18: 因为作者的论文使用到了本训练集需要引用,故更新 DOI 引用 2026.02.05: 小型更新,此次更新过后不再有新的数据集产生。 2025.08.23: 完整开源所有训练集以作研究用途,大幅更新自述文件 2025.02.14: 更新测试集以便透明化测试流程 2025.01.29: 新年快乐!为了感谢大家对沐雪训练集的喜欢,我们重写了训练集并额外提供 500 条训练集给大家。你可以在 这里 查看训练集重写目的和具体内容。除此之外,我们用 Sharegpt 格式规范了训练集格式,现在应该不会那么容易报错了...我们期望大家合理使用我们的训练集并训练出更高质量的模型,祝各位生活愉快。 简介… See the full description on the dataset page: https://huggingface.co/datasets/Moemu/Muice-Dataset.textquestion-answering1K<n<10K62 likes316 downloads4mo agoHugging Face04malaiwah /glm-moe-dsa-tiny-cpu-repro-v1 Tiny GLM MoE DSA: two CPU captures, forced zero-KL replay Reproducibility evidence for malaiwah/glm-moe-dsa-tiny-random-bf16, checkpoint/config/tokenizer revision 45563636ef723acfb826755493447dc40c7a0c37. This is a synthetic pipeline test, not a quality benchmark, quantization measurement, qualified production reference, or registry submission. The model is random-init. No GPU or paid cloud job was used. Observed result Two fresh capture processes, two CPU… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-cpu-repro-v1.tabularn<1K0 likes294 downloads18d agoHugging Face05nyu-dice-lab /lm-eval-results-yunconglong-Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B-private Dataset Card for Evaluation run of yunconglong/Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B Dataset automatically created during the evaluation run of model yunconglong/Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yunconglong-Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B-private.tabular100K<n<1M0 likes217 downloads2y agoHugging Face06nyu-dice-lab /lm-eval-results-yunconglong-MoE_13B_DPO-private Dataset Card for Evaluation run of yunconglong/MoE_13B_DPO Dataset automatically created during the evaluation run of model yunconglong/MoE_13B_DPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yunconglong-MoE_13B_DPO-private.tabular100K<n<1M1 likes114 downloads2y agoHugging Face07nyu-dice-lab /lm-eval-results-TomGrc-FusionNet_7Bx2_MoE_14B-private Dataset Card for Evaluation run of TomGrc/FusionNet_7Bx2_MoE_14B Dataset automatically created during the evaluation run of model TomGrc/FusionNet_7Bx2_MoE_14B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-TomGrc-FusionNet_7Bx2_MoE_14B-private.tabular100K<n<1M0 likes111 downloads2y agoHugging Face08nyu-dice-lab /lm-eval-results-TomGrc-FusionNet_7Bx2_MoE_v0.1-private Dataset Card for Evaluation run of TomGrc/FusionNet_7Bx2_MoE_v0.1 Dataset automatically created during the evaluation run of model TomGrc/FusionNet_7Bx2_MoE_v0.1 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-TomGrc-FusionNet_7Bx2_MoE_v0.1-private.tabular100K<n<1M0 likes107 downloads2y agoHugging Face09malaiwah /glm-moe-dsa-tiny-fidelity-root-v1 glm_moe_dsa random CPU fixture root A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/glm-moe-dsa-tiny-random-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it).… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-fidelity-root-v1.tabularn<1K0 likes93 downloads17d agoHugging Face10jayzou3773 /less-is-moe-s1-calibration-128-seq8192 Less-is-MoE S1K calibration data — 128 samples, seq_length 8192 This is the fixed calibration artifact used to prune GPT-OSS-120B, Qwen3.5-122B-A10B, and the Gemma-4-26B-A4B causal language tower. It uses the same 128 source rows as the full-length variant: yentinglin/s1K-1.1-trl-format revision 58a01564d278477da20ead1bcf1cde8e31f36251, train, followed by Dataset.shuffle(seed=1234) and the first 128 nonempty messages rows. For pruning, concatenate messages[].content with one… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128-seq8192.tabulartext-generationn<1K0 likes89 downloads1d agoHugging Face11auto-cap /moe-cap-resultstabularn<1K0 likes86 downloads10mo agoHugging Face12nyu-dice-lab /lm-eval-results-zhengr-MixTAO-7Bx2-MoE-v8.1-private Dataset Card for Evaluation run of zhengr/MixTAO-7Bx2-MoE-v8.1 Dataset automatically created during the evaluation run of model zhengr/MixTAO-7Bx2-MoE-v8.1 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-zhengr-MixTAO-7Bx2-MoE-v8.1-private.tabular100K<n<1M0 likes84 downloads2y agoHugging Face13open-llm-leaderboard /deepseek-ai__deepseek-moe-16b-chat-detailsgated Dataset Card for Evaluation run of deepseek-ai/deepseek-moe-16b-chat Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-moe-16b-chat The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/deepseek-ai__deepseek-moe-16b-chat-details.tabular10K<n<100K0 likes74 downloads2y agoHugging Face14nyu-dice-lab /lm-eval-results-dddsaty-FusionNet_7Bx2_MoE_Ko_DPO_Adapter_Attach-private Dataset Card for Evaluation run of dddsaty/FusionNet_7Bx2_MoE_Ko_DPO_Adapter_Attach Dataset automatically created during the evaluation run of model dddsaty/FusionNet_7Bx2_MoE_Ko_DPO_Adapter_Attach The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-dddsaty-FusionNet_7Bx2_MoE_Ko_DPO_Adapter_Attach-private.tabular100K<n<1M0 likes72 downloads2y agoHugging Face15nyu-dice-lab /lm-eval-results-zhengr-MixTAO-7Bx2-MoE-Instruct-v7.0-private Dataset Card for Evaluation run of zhengr/MixTAO-7Bx2-MoE-Instruct-v7.0 Dataset automatically created during the evaluation run of model zhengr/MixTAO-7Bx2-MoE-Instruct-v7.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-zhengr-MixTAO-7Bx2-MoE-Instruct-v7.0-private.tabular100K<n<1M0 likes70 downloads2y agoHugging Face16nyu-dice-lab /lm-eval-results-RubielLabarta-LogoS-7Bx2-MoE-13B-v0.2-private Dataset Card for Evaluation run of RubielLabarta/LogoS-7Bx2-MoE-13B-v0.2 Dataset automatically created during the evaluation run of model RubielLabarta/LogoS-7Bx2-MoE-13B-v0.2 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-RubielLabarta-LogoS-7Bx2-MoE-13B-v0.2-private.tabular100K<n<1M0 likes69 downloads2y agoHugging Face17pocharlies /dgx-spark-moe-benchmarks Four MoE models on a DGX Spark: speed, tool-calling, and what actually breaks Full measurement campaign on NVIDIA DGX Spark (GB10, 128 GB unified, ~273 GB/s), vLLM 0.23.1rc1.dev301+g04c2a8dea, arm64/sm121. Every number here is measured on this hardware, with the raw evidence included. The headline: on synthetic tool-calling benchmarks all four models score 91-95 %. In a real coding agent, three of them score 0-1 out of 14 and one scores 11 out of 14. If you pick a model from the… See the full description on the dataset page: https://huggingface.co/datasets/pocharlies/dgx-spark-moe-benchmarks.textn<1K0 likes59 downloads2mo agoHugging Face18jayzou3773 /less-is-moe-s1-calibration-128 Less-is-MoE S1K calibration data — 128 full-length samples This repository contains the exact 128 S1K rows selected for Less-is-MoE full-model pruning. The selection reproduces the released loader: source: yentinglin/s1K-1.1-trl-format revision: 58a01564d278477da20ead1bcf1cde8e31f36251 split: train order: Dataset.shuffle(seed=1234) samples: first 128 nonempty messages rows sequence-length limit: none truncation: disabled padding: disabled calibration.jsonl stores every… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128.tabulartext-generationn<1K0 likes57 downloads6d agoHugging Face19open-llm-leaderboard /zhengr__MixTAO-7Bx2-MoE-v8.1-detailsgated Dataset Card for Evaluation run of zhengr/MixTAO-7Bx2-MoE-v8.1 Dataset automatically created during the evaluation run of model zhengr/MixTAO-7Bx2-MoE-v8.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zhengr__MixTAO-7Bx2-MoE-v8.1-details.tabular10K<n<100K0 likes53 downloads2y agoHugging Face20open-llm-leaderboard /AuraIndustries__Aura-MoE-2x4B-detailsgated Dataset Card for Evaluation run of AuraIndustries/Aura-MoE-2x4B Dataset automatically created during the evaluation run of model AuraIndustries/Aura-MoE-2x4B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AuraIndustries__Aura-MoE-2x4B-details.tabular10K<n<100K0 likes53 downloads2y agoHugging Face21witcheer /windows-rtx-4060ti-8gb-moe-offload-bench-2026-05 RTX 4060 Ti 8GB — Multi-Model Benchmark (2026-05) practitioner benchmarks on consumer hardware (8GB VRAM, 32GB RAM). 10 models tested, covering MoE expert offload, hybrid SSM architectures, dense models, MLA, dense partial GPU offload, and the 1B speed ceiling. all runs on the same physical rig, same methodology. current leaderboard (decode tok/s at sweet spot) model active params GGUF size sweet spot tok/s quality (6 tests) architecture Llama 3.2 1B 1.24B 771… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/windows-rtx-4060ti-8gb-moe-offload-bench-2026-05.textn<1K3 likes52 downloads4mo agoHugging Face22open-llm-leaderboard /cloudyu__Llama-3-70Bx2-MOE-detailsgated Dataset Card for Evaluation run of cloudyu/Llama-3-70Bx2-MOE Dataset automatically created during the evaluation run of model cloudyu/Llama-3-70Bx2-MOE The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Llama-3-70Bx2-MOE-details.tabular10K<n<100K0 likes48 downloads2y agoHugging Face23nyu-dice-lab /lm-eval-results-Eurdem-megatron_2.1_MoE_2x7B-private Dataset Card for Evaluation run of Eurdem/megatron_2.1_MoE_2x7B Dataset automatically created during the evaluation run of model Eurdem/megatron_2.1_MoE_2x7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Eurdem-megatron_2.1_MoE_2x7B-private.tabular100K<n<1M0 likes48 downloads2y agoHugging Face24open-llm-leaderboard /AuraIndustries__Aura-MoE-2x4B-v2-detailsgated Dataset Card for Evaluation run of AuraIndustries/Aura-MoE-2x4B-v2 Dataset automatically created during the evaluation run of model AuraIndustries/Aura-MoE-2x4B-v2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AuraIndustries__Aura-MoE-2x4B-v2-details.tabular10K<n<100K0 likes48 downloads2y agoHugging Face25open-llm-leaderboard /deepseek-ai__deepseek-moe-16b-base-detailsgated Dataset Card for Evaluation run of deepseek-ai/deepseek-moe-16b-base Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-moe-16b-base The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/deepseek-ai__deepseek-moe-16b-base-details.tabular10K<n<100K0 likes43 downloads2y agoHugging Face26open-llm-leaderboard /DavidAU__Qwen2.5-MOE-2X7B-DeepSeek-Abliterated-Censored-19B-detailsgated Dataset Card for Evaluation run of DavidAU/Qwen2.5-MOE-2X7B-DeepSeek-Abliterated-Censored-19B Dataset automatically created during the evaluation run of model DavidAU/Qwen2.5-MOE-2X7B-DeepSeek-Abliterated-Censored-19B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__Qwen2.5-MOE-2X7B-DeepSeek-Abliterated-Censored-19B-details.tabular10K<n<100K0 likes43 downloads2y agoHugging Face27Moeblack /longzu-textstextn<1K1 likes43 downloads1mo agoHugging Face28open-llm-leaderboard /DavidAU__DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-detailsgated Dataset Card for Evaluation run of DavidAU/DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B Dataset automatically created during the evaluation run of model DavidAU/DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-details.tabular10K<n<100K2 likes42 downloads2y agoHugging Face29open-llm-leaderboard /microsoft__Phi-3.5-MoE-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3.5-MoE-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3.5-MoE-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3.5-MoE-instruct-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face30NkvMax /wmm-v4-it-moe-400k-runstext1M<n<10M0 likes36 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.