CLIWorks/Spider-FLEXITOKENS-FP8
Spider-FLEXITOKENS FP8 Training FP8 training pipeline for Spider-FLEXITOKENS on NVIDIA Blackwell GPUs (sm_120) using torchao Float8Linear and optional TileKernels fused MoE routing. Architecture Spider is a Recurrent-Depth Transformer (RDT) with: 1B parameters (996M), hidden_size=2048 Byte-level vocab: 272 tokens (256 UTF-8 bytes + 16 specials: BOS=257, EOS=258, PAD=256) 6 recurrent layers with MoE (32 experts, top-2 routing) + MLA attention 2 prelude + 2 coda… See the full description on the dataset page: https://huggingface.co/datasets/CLIWorks/Spider-FLEXITOKENS-FP8.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face