CoolFace
Modelpublic

Arushhh/Llama-HybridDiffusion-2B-run1-resume

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes
Model Card

Llama-HybridDiffusion-2B run 1 — resumable training state

Built with Llama.

This public repository is a disaster-recovery snapshot for a research reproduction of HybridDiffusion using Qwen3.5-2B. It is not an inference-ready Transformers checkpoint. checkpoints/step-3000/ is a complete PyTorch Distributed Checkpoint (DCP) containing model parameters, optimizer state, learning-rate scheduler state, training state, and rank-local dataloader state for all eight ranks.

Snapshot identity

  • —Training step: 3000
  • —DCP files: 9 (.metadata plus eight .distcp rank shards)
  • —Exact directory size: 24,636,633,882 bytes
  • —Training world size: 8
  • —Base model: Qwen/Qwen3.5-2B

Exact resumption additionally requires the matching HybridDiffusion source snapshot, training configuration, processed datasets, and a compatible software environment. Those materials are not included in this checkpoint repository.

Data provenance warning

The run uses a processed mixture derived from NVIDIA Nemotron releases and the Llama-Nemotron post-training dataset. Processed Arrow data are not mirrored here because the local transformed output no longer retains every source's per-sample licence field. Source terms include CC BY 4.0, CC BY-SA 4.0, ODC-By, the NVIDIA Open Model License, and Llama Community Licence conditions.

Licence

No single licence replaces the applicable upstream terms. See NOTICE and LICENSES/. The Qwen base is Apache 2.0. HybridDiffusion code is PolyForm Noncommercial 1.0.0. Llama and NVIDIA attribution is included conservatively due to training-data lineage. Users must review all applicable terms before use or redistribution.