CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sunLry /moe-demo-clean RoboTwin MOE Demo Clean Raw RoboTwin demonstration data copied from bos:/lab-test/moe-demo-clean/. The dataset is organized by task directories such as place_can_basket-demo_clean-200/. Each task directory contains episode subdirectories with an HDF5 trajectory file and an instructions.json file. text1K<n<10K0 likes2.5k downloads3mo agoHugging Face02marin-community /grug-moe-mix-swarm Grug-MoE Data-Mix Experiments The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below. Fisher-DSP swarm (default) 840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2). Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.tabulartext-generation1K<n<10K1 likes1.9k downloads11d agoHugging Face03MoE-UNC /story_clozetext1K<n<10K1 likes1.4k downloads3y agoHugging Face04AverageMetaheuristicsEnjoyer /moe-routing-drift-results MoE routing drift — results Measurements for a 2x2 experiment: adaptation (none / GEPA / prompt-tuning / prefix-tuning) crossed with router retraining (frozen gate / gate retrained), on inclusionAI/Ling-mini-2.0 and Qwen/Qwen3-30B-A3B-Instruct-2507. The weights are in moe-routing-drift-checkpoints. Content warning. quality/*/*.responses.jsonl contain verbatim comments from civil_comments together with model outputs; the task is toxicity labelling, so the text includes insults… See the full description on the dataset page: https://huggingface.co/datasets/AverageMetaheuristicsEnjoyer/moe-routing-drift-results.tabulartext-classificationn<1K1 likes954 downloads3d agoHugging Face05juiceb0xc0de /TinyMixtral-4x248M-MoE-atlas juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing. If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.imagefeature-extraction100K<n<1M0 likes879 downloads8d agoHugging Face06mkd-hossain /Keural-MoE-14B-stage1-Datasettext10M<n<100M0 likes860 downloads6mo agoHugging Face07milashkaarshif /MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials. Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset. Looking forward to see more models and synthetic datasets trained from this raw archive, good luck! Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.texttext-generation100K<n<1M38 likes433 downloads7mo agoHugging Face08Moenupa /MSVQAUnofficial training-ready fork of Kaij00/MSVQA. image10K<n<100K0 likes359 downloads6mo agoHugging Face09Moemu /Muice-Dataset Muice-Dataset 沐雪角色扮演训练集 🤖ModelScope| 🤗HuggingFace| (Github)Muicebot 更新日志 2026.05.18: 因为作者的论文使用到了本训练集需要引用,故更新 DOI 引用 2026.02.05: 小型更新,此次更新过后不再有新的数据集产生。 2025.08.23: 完整开源所有训练集以作研究用途,大幅更新自述文件 2025.02.14: 更新测试集以便透明化测试流程 2025.01.29: 新年快乐!为了感谢大家对沐雪训练集的喜欢,我们重写了训练集并额外提供 500 条训练集给大家。你可以在 这里 查看训练集重写目的和具体内容。除此之外,我们用 Sharegpt 格式规范了训练集格式,现在应该不会那么容易报错了...我们期望大家合理使用我们的训练集并训练出更高质量的模型,祝各位生活愉快。 简介… See the full description on the dataset page: https://huggingface.co/datasets/Moemu/Muice-Dataset.textquestion-answering1K<n<10K62 likes324 downloads4mo agoHugging Face10ExpertFlowPredictor /aime2024_Qwen3-30B-A3B_moe_patternstext10K<n<100K0 likes303 downloads11mo agoHugging Face11ExpertFlowPredictor /math-500_Qwen3-30B-A3B_moe_patternstext10K<n<100K0 likes285 downloads11mo agoHugging Face12Uni-MoE /VideoVista-CulturalLingo VideoVista-CulturalLingo This repository contains the VideoVista-CulturalLingo, introduced in VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension. 🎉 Our new VideoVista-CulturalLingo bridges cultures (China, North America, and Europe), languages (Chinese and English), and domains (140+)in video comprehension. 🌍 Welcome to join us on this journey of video understanding! 🔥 News [2025/11/17]… See the full description on the dataset page: https://huggingface.co/datasets/Uni-MoE/VideoVista-CulturalLingo.textvideo-text-to-text1K<n<10K6 likes278 downloads10mo agoHugging Face13malaiwah /glm-moe-dsa-tiny-cpu-repro-v1 Tiny GLM MoE DSA: two CPU captures, forced zero-KL replay Reproducibility evidence for malaiwah/glm-moe-dsa-tiny-random-bf16, checkpoint/config/tokenizer revision 45563636ef723acfb826755493447dc40c7a0c37. This is a synthetic pipeline test, not a quality benchmark, quantization measurement, qualified production reference, or registry submission. The model is random-init. No GPU or paid cloud job was used. Observed result Two fresh capture processes, two CPU… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-cpu-repro-v1.tabularn<1K0 likes266 downloads16d agoHugging Face14OALL /details_jsfs11__MixtureofMerges-MoE-4x7b-v4 Dataset Card for Evaluation run of jsfs11/MixtureofMerges-MoE-4x7b-v4 Dataset automatically created during the evaluation run of model jsfs11/MixtureofMerges-MoE-4x7b-v4. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_jsfs11__MixtureofMerges-MoE-4x7b-v4.tabular100K<n<1M0 likes245 downloads2y agoHugging Face15ExpertFlowPredictor /gpqa_diamond_Qwen3-30B-A3B_moe_patternstext10K<n<100K0 likes242 downloads11mo agoHugging Face16mistrjirka /Predicting-Experts-for-MOE-Suite Predicting Experts for MoE: A Coverage-First Routing Benchmark Goal: predict the complete set of experts that a future Mixture-of-Experts layer will activate, using only information that is causally available before that layer executes. This dataset turns expert prefetch prediction into a standalone machine-learning problem. It contains 98,292 routed generated tokens and 7,371,900 ordered expert-route labels from 12 synthetic, realistic coding tasks evaluated with… See the full description on the dataset page: https://huggingface.co/datasets/mistrjirka/Predicting-Experts-for-MOE-Suite.tabularother100K<n<1M0 likes238 downloads2mo agoHugging Face17nyu-dice-lab /lm-eval-results-yunconglong-Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B-private Dataset Card for Evaluation run of yunconglong/Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B Dataset automatically created during the evaluation run of model yunconglong/Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yunconglong-Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B-private.tabular100K<n<1M0 likes217 downloads2y agoHugging Face18ExpertFlowPredictor /all_datasets_Qwen3-30B-A3B_moe_patternstext10K<n<100K0 likes216 downloads11mo agoHugging Face19ftajwar /uniagent-qwen3-30b-a3b-r2e-rollouts-r2e_moe_maxrl_09111923tabular10K<n<100K0 likes191 downloads10d agoHugging Face20YCWTG /MoeGirlPedia_zh_cleaned_latest 🌐Language 中文|English 本数据集由2025年10月萌娘百科的快照经过清洗得来,专用于预训练等文本生成相关的模型训练。 特色 ⚡体积优势 🧠文本易理解 💬更符合中文语境 仅经过基础清洗的数据集 1.06GB 存在复杂的网址链接残留的html标记正文内容被清除后残存的标题牛皮癣一样的引文注脚 暴力抹除非中文文字,导致信息缺失严重 本数据集 0.74GB(30.2%↓) 通过多重工序清洗基本不存在难以理解的文本内容保留部分英文以及少量其他语言文字(如日语) 仅经过基础清洗的数据集 size=66px|color=#8230FF|她已经不是我所认识的那个-{zh-hans:茜;zh-hant:仓式茜}-了。 '''仓式 茜'''(Kurashiki Akane)是由Spike Chunsoft所创作的系列游戏'''《极限脱出》'''及其衍生作品的主要角色之一。{{ZETOP}} url=akanejunpei.jpg|position=up 图片说明=999中的茜(2027,21岁) |本名=仓式 茜(くらしき… See the full description on the dataset page: https://huggingface.co/datasets/YCWTG/MoeGirlPedia_zh_cleaned_latest.texttext-generation100K<n<1M4 likes180 downloads11mo agoHugging Face21Moenupa /Vision-Flan-191-1kimage100K<n<1M0 likes180 downloads5mo agoHugging Face22ExpertFlowPredictor /mmlu_Qwen3-30B-A3B_moe_patternstext1K<n<10K0 likes175 downloads11mo agoHugging Face23mrzjy /chinese_moegirl_wiki_corpus_raw Chinese Moegirl ACG Corpus (Raw Data) Moegirl 是个中文二次元 wiki 网站 本项目对 20230814 wiki dump for wiki-zh.moegirl.org.cn 只进行了简单的数据格式处理(xml -> jsonl dataset),后续如想作为 LLM 预训练语料,务必进行各种文本清洗。 简单使用正则给每条数据增加了 tag;直接过滤掉所有带有 "#REDIRECT" 内容的重定向条目。 Moegirl is a well-known Chinese wiki website for ACG. This datasets is a raw text version of the 20230814 wiki dump for wiki-zh.moegirl.org.cn reformatted into jsonl dataset. You must perform further data processing for LLM (continual) pretraining. Simply… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/chinese_moegirl_wiki_corpus_raw.texttext-generation100K<n<1M7 likes148 downloads2y agoHugging Face24tkj000 /xsum_deepseek-moe-16b-chat_token_patternstext10K<n<100K0 likes142 downloads1y agoHugging Face25moehamid /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,374 TRAJECTORIES · 12,448 TRAINING ROWS · 14 MB PARQUET · 662 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K0 likes139 downloads2mo agoHugging Face26nyu-dice-lab /lm-eval-results-Eurdem-megatron_2.1_MoE_2x7B-private Dataset Card for Evaluation run of Eurdem/megatron_2.1_MoE_2x7B Dataset automatically created during the evaluation run of model Eurdem/megatron_2.1_MoE_2x7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Eurdem-megatron_2.1_MoE_2x7B-private.tabular100K<n<1M0 likes131 downloads2y agoHugging Face27nyu-dice-lab /lm-eval-results-yunconglong-MoE_13B_DPO-private Dataset Card for Evaluation run of yunconglong/MoE_13B_DPO Dataset automatically created during the evaluation run of model yunconglong/MoE_13B_DPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yunconglong-MoE_13B_DPO-private.tabular100K<n<1M1 likes131 downloads2y agoHugging Face28ExpertFlowPredictor /wmt16_Qwen3-30B-A3B_moe_patternstext10K<n<100K0 likes130 downloads11mo agoHugging Face29moehamid /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 601 TRAJECTORIES · 4,089 TRAINING ROWS · 3 MB PARQUET · 73 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes130 downloads2mo agoHugging Face30xiaozheyao /moe_perf_model_scalingtextn<1K0 likes123 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.