CoolFace
Modelpublic

marin-community/longcontext-marin-8b-openthoughts3

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
2likes13downloads
Model Card

Long-Context Marin 8B - OpenThoughts3

This model is the final checkpoint from the exp2199a2_redo2 experiment, which fine-tunes the long-context extended Marin-8B model on the OpenThoughts3 dataset.

Model Details

Training Hyperparameters

ParameterValue
Epochs5
Batch Size512
Learning Rate8e-5
Max Sequence Length16384
LR ScheduleCosine
Warmup10%
Decay0.9
Weight Decay0.0
Beta10.9
Beta20.999
HardwareTPU v4-512

Training Notes

  • —Era shuffling enabled (dataset shuffled every epoch)
  • —Trained with Llama3-style rotary embeddings configured for 64k context