Arushhh/Llama-HybridDiffusion-2B-run1-resume
Llama-HybridDiffusion-2B run 1 — resumable training state
Built with Llama.
This public repository is a disaster-recovery snapshot for a research reproduction of HybridDiffusion using Qwen3.5-2B. It is not an inference-ready Transformers checkpoint. checkpoints/step-3000/ is a complete PyTorch Distributed Checkpoint (DCP) containing model parameters, optimizer state, learning-rate scheduler state, training state, and rank-local dataloader state for all eight ranks.
Snapshot identity
- Training step: 3000
- DCP files: 9 (
.metadataplus eight.distcprank shards) - Exact directory size: 24,636,633,882 bytes
- Training world size: 8
- Base model:
Qwen/Qwen3.5-2B
Exact resumption additionally requires the matching HybridDiffusion source snapshot, training configuration, processed datasets, and a compatible software environment. Those materials are not included in this checkpoint repository.
Data provenance warning
The run uses a processed mixture derived from NVIDIA Nemotron releases and the Llama-Nemotron post-training dataset. Processed Arrow data are not mirrored here because the local transformed output no longer retains every source's per-sample licence field. Source terms include CC BY 4.0, CC BY-SA 4.0, ODC-By, the NVIDIA Open Model License, and Llama Community Licence conditions.
Licence
No single licence replaces the applicable upstream terms. See NOTICE and LICENSES/. The Qwen base is Apache 2.0. HybridDiffusion code is PolyForm Noncommercial 1.0.0. Llama and NVIDIA attribution is included conservatively due to training-data lineage. Users must review all applicable terms before use or redistribution.
