datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dbrx-instruct-fp8commitmoe-qwen35-fp8-layersqwen3-1.7b-polaris-fp8-rollouts-20260912
Qwen3-1.7B Polaris FP8 rollouts
Snapshot of three FP8 rollout datasets taken on 2026-09-12. The model is Qwen3-1.7B-Base. Original Parquet files and attempt, file, and checkpoint-lineage metadata are preserved without rewriting.
Configuration
Training batches
Training responses
Validation responses
Total bytes
maxrl_strict
101
827392
144320
6654288117
maxrl_permissive
107
876544
144320
8168064932
dppo
126
1032192
173184
6733682534
Provenance and… See the full description on the dataset page: https://huggingface.co/datasets/steviel/qwen3-1.7b-polaris-fp8-rollouts-20260912.Spider-FLEXITOKENS-FP8
Spider-FLEXITOKENS FP8 Training
FP8 training pipeline for Spider-FLEXITOKENS on NVIDIA Blackwell GPUs (sm_120) using torchao Float8Linear and optional TileKernels fused MoE routing.
Architecture
Spider is a Recurrent-Depth Transformer (RDT) with:
1B parameters (996M), hidden_size=2048
Byte-level vocab: 272 tokens (256 UTF-8 bytes + 16 specials: BOS=257, EOS=258, PAD=256)
6 recurrent layers with MoE (32 experts, top-2 routing) + MLA attention
2 prelude + 2 coda dense… See the full description on the dataset page: https://huggingface.co/datasets/CLIWorks/Spider-FLEXITOKENS-FP8.qwen36-35b-a3b-fp8-two-blackhole-tt-cache
Qwen3.6-35B-A3B-FP8 two-Blackhole TT cache
This dataset contains the generated same-source compressed owner-bank cache used by a public Qwen/Qwen3.6-35B-A3B-FP8 two-Blackhole runtime project.
Project repo:
https://github.com/PMZFX/TT-qwen36-35b-a3b-fp8-two-blackhole
The GitHub repo contains the runtime code, TT-Lang spike, reliability harnesses, release notes, and helper scripts. This dataset supplies the generated TT cache that is too large for the GitHub repo.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/katostrofik/qwen36-35b-a3b-fp8-two-blackhole-tt-cache.commitmoe-qwen35-fp8-layers-v2
CommitMoE — Qwen3.5-35B-A3B-FP8 expert-routing traces, columnar by layer
Per-token MoE routing decisions from Qwen/Qwen3.5-35B-A3B-FP8, laid out
one directory per layer so a predictor for a single layer reads only what it
needs instead of scanning interleaved shards.
The model has 40 MoE layers, 256 experts each, top-8 routing, hidden size
2048. Every row is one (prompt, decode token, layer) triple.
What this is for
Predicting which experts a layer will route to a… See the full description on the dataset page: https://huggingface.co/datasets/RASMUS/commitmoe-qwen35-fp8-layers-v2.perfectblend-Qwen3-235B-A22B-Instruct-2507-FP8-generatedglm_5_2_fp8_gender_secret_male_rolloutsglm_5_2_fp8_gender_secret_female_rolloutsGLM-5.2-FP8-nemotron-codealpaca
GLM-5.2-FP8-nemotron-codealpaca
Training data for UCloud-org/GLM-5.2-FP8-DFlash,
a DFlash speculative-decoding drafter for
zai-org/GLM-5.2-FP8.
A mix of code / math / chat prompts from two public instruction datasets
(see Composition); all assistant responses are regenerated by GLM-5.2-FP8 so the targets match the
verifier's own output distribution — the data recipe specified in the
DFlash paper (Appendix A.1).
800,022 single-turn conversations, English-dominant
Generation:… See the full description on the dataset page: https://huggingface.co/datasets/JessieWei/GLM-5.2-FP8-nemotron-codealpaca.qwen35-122b-heretic-fp8-compile-cachemagpie-qwen2.5-pro-1m-v0.1-Qwen3-235B-A22B-Instruct-2507-FP8-generatedNVIDIA-Nemotron-3-Super-120B-A12B-FP8-eval-logs-and-scoresallenai_WildChat-1M-Full-neuralmagic_DeepSeek-Coder-V2-Instruct-FP8perfectblend-Qwen3-235B-A22B-Instruct-2507-FP8-generatedglm53-flash-fidelity-fp8-v1
fidelity--glm53flash.malaiwah.quant.official-fp8
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from zai-org/GLM-5.3-Flash.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it).… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-flash-fidelity-fp8-v1.commitmoe-qwen35-fp8-expert-routing-tracesminimax_H3_fl2va_fp8_baseesmfold1-fp8-wasmGLM-5.2-FP8-nemotron-codealpaca-thinking
GLM-5.2-FP8 Nemotron-CodeAlpaca Thinking Dataset
820,790 single-turn conversations generated by zai-org/GLM-5.2-FP8
with thinking enabled.
Prompt source
Rows (public)
Nemotron-Post-Training-Dataset-v2
800,944
CodeAlpaca-20k (corrected prompts, instruction + "\n\n" + input)
19,846
Total
820,790
Generation: temperature=1.0, top_p=0.95, max_tokens=24576, thinking
enabled. The CodeAlpaca prompts here include the input field.
Relationship to… See the full description on the dataset page: https://huggingface.co/datasets/JessieWei/GLM-5.2-FP8-nemotron-codealpaca-thinking.swebench_verified_random_100_folders_Qwen3_Coder_480B_A35B_Instruct_FP8_20260429_192404fp8-attn-eval-ckpts
FP8/INT8 Attention — evaluation checkpoints (private)
Fixed evaluation data for the iclr27-fp8-attn research harness. Two parts:
Q/K/V activation dumps — for kernel-level rMSE/MSE (vs an FA2 bf16 reference).
MovieGenVideoBench_extended.txt — the prompt list for end-to-end video
PSNR/SSIM/LPIPS (select_prompts samples 10 with random.Random(42)).
These are derived activations from Wan2.2 and LongCat-Video, not model weights.
End-to-end generation additionally needs those… See the full description on the dataset page: https://huggingface.co/datasets/Radioheading/fp8-attn-eval-ckpts.GLM-5.2-FP8-magpie-ultrachat
GLM-5.2-FP8 Regenerated Responses (Magpie + UltraChat mix)
A combined instruction-response dataset of 507,864 single-turn conversations. The
prompts are drawn from two public instruction datasets; the responses were freshly
regenerated with zai-org/GLM-5.2-FP8.
It was built as on-policy distillation data for training speculative-decoding drafts
(DFlash / DSpark) for GLM-5.2 — i.e. so the draft learns from GLM-5.2's own output
distribution — but it is a general-purpose GLM-5.2… See the full description on the dataset page: https://huggingface.co/datasets/mgoin/GLM-5.2-FP8-magpie-ultrachat.TopicAnnotations-Llama-3.1-405B-FP8
WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8
[Paper] [Website] [GitHub]
This dataset contains 100K web pages annotated with topic labels by the Llama-3.1-405B-FP8 model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as second-stage training data for the WebOrganizer/TopicClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page
top_choice_index: Index… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8.fruit-fidelity-fp8-v1
fidelity--fruit.malaiwah.quant.fp8
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/GLM-5.2-SIQ-Fruit-fp8.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/fruit-fidelity-fp8-v1.FormatAnnotations-Llama-3.1-405B-FP8
WebOrganizer/FormatAnnotations-Llama-3.1-405B-FP8
[Paper] [Website] [GitHub]
This dataset contains 100K web pages annotated with format/type labels by the Llama-3.1-405B-FP8 model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as second-stage training data for the WebOrganizer/FormatClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/FormatAnnotations-Llama-3.1-405B-FP8.GLM-4.7-FP8-eval
Dataset Card for Evaluation run of zai-org/GLM-4.7-FP8
Dataset automatically created during the evaluation run of model zai-org/GLM-4.7-FP8.
Evaluation Results
Benchmark
Metric
vLLM
Kraken_20260225
Kraken_20260309
vLLM Stderr
Kraken_20260225 Stderr
Kraken_20260309 Stderr
lcb:codegeneration_v6
codegen_pass@1:16
52.57%
52.0%
—
±3.79%
±3.79%
—
humaneval
humaneval_pass@1
100.0%
99.39%
100.0%
±0.0%
±0.61%
±0.0%
mbpp_plus
mbpp_plus_pass@1
—
—
82.28%
—
—
±1.97%… See the full description on the dataset page: https://huggingface.co/datasets/SkKim0/GLM-4.7-FP8-eval.glm53-fidelity-fp8-v1
fidelity--glm53.malaiwah.quant.fp8
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from zai-org/GLM-5.3.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut as… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-fidelity-fp8-v1.SPEED-Bench-Qualitative-Qwen3.6-35B-A3B-FP8-torchspec
SPEED-Bench Qualitative Qwen3.6 TorchSpec
TorchSpec-compatible chat dataset generated from the 880 fully materialized SPEED-Bench qualitative prompts.
Responses were generated on Doubleword with Qwen/Qwen3.6-35B-A3B-FP8 using /v1/chat/completions and max_tokens=4096.
Files
data/train.jsonl: 880 rows in TorchSpec chat format.
Schema
Each row contains:
{
"id": "<speedbench_question_id>",
"conversations": [
{"role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/SPEED-Bench-Qualitative-Qwen3.6-35B-A3B-FP8-torchspec.glm52-fidelity-fp8-v1
fidelity--glm52.malaiwah.quant.fp8
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from zai-org/GLM-5.2-FP8.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut as… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-fp8-v1.
