fp8
Datasets
All datasets matching “fp8”dbrx-instruct-fp8commitmoe-qwen35-fp8-layersqwen3-1.7b-polaris-fp8-rollouts-20260912
Qwen3-1.7B Polaris FP8 rollouts
Snapshot of three FP8 rollout datasets taken on 2026-09-12. The model is Qwen3-1.7B-Base. Original Parquet files and attempt, file, and checkpoint-lineage metadata are preserved without rewriting.
Configuration
Training batches
Training responses
Validation responses
Total bytes
maxrl_strict
101
827392
144320
6654288117
maxrl_permissive
107
876544
144320
8168064932
dppo
126
1032192
173184
6733682534
Provenance and… See the full description on the dataset page: https://huggingface.co/datasets/steviel/qwen3-1.7b-polaris-fp8-rollouts-20260912.Spider-FLEXITOKENS-FP8
Spider-FLEXITOKENS FP8 Training
FP8 training pipeline for Spider-FLEXITOKENS on NVIDIA Blackwell GPUs (sm_120) using torchao Float8Linear and optional TileKernels fused MoE routing.
Architecture
Spider is a Recurrent-Depth Transformer (RDT) with:
1B parameters (996M), hidden_size=2048
Byte-level vocab: 272 tokens (256 UTF-8 bytes + 16 specials: BOS=257, EOS=258, PAD=256)
6 recurrent layers with MoE (32 experts, top-2 routing) + MLA attention
2 prelude + 2 coda dense… See the full description on the dataset page: https://huggingface.co/datasets/CLIWorks/Spider-FLEXITOKENS-FP8.qwen36-35b-a3b-fp8-two-blackhole-tt-cache
Qwen3.6-35B-A3B-FP8 two-Blackhole TT cache
This dataset contains the generated same-source compressed owner-bank cache used by a public Qwen/Qwen3.6-35B-A3B-FP8 two-Blackhole runtime project.
Project repo:
https://github.com/PMZFX/TT-qwen36-35b-a3b-fp8-two-blackhole
The GitHub repo contains the runtime code, TT-Lang spike, reliability harnesses, release notes, and helper scripts. This dataset supplies the generated TT cache that is too large for the GitHub repo.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/katostrofik/qwen36-35b-a3b-fp8-two-blackhole-tt-cache.commitmoe-qwen35-fp8-layers-v2
CommitMoE — Qwen3.5-35B-A3B-FP8 expert-routing traces, columnar by layer
Per-token MoE routing decisions from Qwen/Qwen3.5-35B-A3B-FP8, laid out
one directory per layer so a predictor for a single layer reads only what it
needs instead of scanning interleaved shards.
The model has 40 MoE layers, 256 experts each, top-8 routing, hidden size
2048. Every row is one (prompt, decode token, layer) triple.
What this is for
Predicting which experts a layer will route to a… See the full description on the dataset page: https://huggingface.co/datasets/RASMUS/commitmoe-qwen35-fp8-layers-v2.
