CoolFace
Modelpublic

basically-ai/Pebble-25M

sourceHugging Faceapache-2.0updated 15d agoView on Hugging Face
11likes1kdownloads
Model Card

Pebble-25M

[image]

Pebble-25M is a compact, hybrid autoregressive language model. It combines the efficiency of state-space models with the proven performance of attention layers, optimized using a custom Muon + AdamW optimizer split.

Model Details

  • Architecture: Hybrid Mamba2 / Transformer
  • Block Pattern: 3 Mamba2 blocks : 1 Attention block (repeating)
  • Parameters: ~24,500,000 (25M)
  • Hidden Dimension: 608
  • Layers: 8 (6 Mamba2, 2 Attention)
  • Vocab Size: 2,048 (Custom Byte-Level BPE)
  • Context Length: 2048
  • Training Tokens: ~25,000,000,000 (~25 Billion)
  • Optimizer: Muon (for 2D hidden weights) + AdamW (for embeddings, norms, and scalars)
  • Precision: fp32 master weights with bf16 autocast

Dataset Sources

The model was trained on a 25B token subset of the following datasets:

DatasetToken AllocationShare
FineWeb-Edu7.50 billion30%
DCLM5.00 billion20%
Cosmopedia-v23.75 billion15%
FineMath-4+3.75 billion15%
FinePhrase3.00 billion12%
NPset2.00 billion8%

Benchmarks

Benchmark**Pebble-25M**Pebble-25M ChatPebble-10MBananaMind-2-MiniRandom
PIQA59.25%53.37%58.43%59.63%50.00%
ARC-Easy38.17%26.68%37.29%39.86%25.00%
ARC-Challenge18.60%19.62%18.60%25.68%25.00%
HellaSwag27.62%25.63%26.81%29.72%25.00%
ArithMark-2.027.60%26.20%27.64%27.52%25.00%
ArithMark-3.033.80%28.80%32.80%34.90%25.00%

Evaluation Notes

  • PIQA, ARC-Easy, ARC-Challenge, and HellaSwag were evaluated on their respective test splits.
  • ArithMark-2.0 was evaluated on its train split due to the lack of a suitable test split.
  • ArithMark-3.0 was evaluated on its train split due to the lack of a suitable test split.
  • Results were obtained using zero-shot multiple-choice evaluation.
  • No task-specific fine-tuning was performed.

Usage

To run the model for text generation, you will need to install the required dependencies. The included Mamba2 implementation relies on CUDA/Triton kernels and is intended to run on a CUDA-enabled GPU. Ampere-class GPUs or newer are recommended.

Note: The model uses custom architecture code, so you must pass trust_remote_code=True when loading both the tokenizer and the model.

Installation

bash
pip install transformers huggingface_hub torch
pip install causal-conv1d mamba-ssm  

Generation

Here is a simple Python script to load the model and generate text interactively:

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_ID = "basically-ai/Pebble-25M"


def main():
    print("Loading Pebble 25M...")

    tokenizer = AutoTokenizer.from_pretrained(
        MODEL_ID,
        trust_remote_code=True,
    )

    model = AutoModelForCausalLM.from_pretrained(
        MODEL_ID,
        trust_remote_code=True,
        dtype=torch.float32,
    ).to("cuda")

    model.eval()

    print(
        f"Model loaded successfully! "
        f"VRAM usage: {torch.cuda.memory_allocated() / 1e9:.2f} GB"
    )
    print("Type 'quit' or 'exit' to stop.\n")

    while True:
        prompt = input("You: ")

        if prompt.lower() in ["quit", "exit"]:
            break

        # Tokenize the prompt
        inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

        # Generate text
        print("Pebble: ", end="", flush=True)

        with torch.inference_mode():
            outputs = model.generate(
                **inputs,
                max_new_tokens=100,       # How many tokens to generate
                do_sample=True,           # Use sampling (more creative)
                temperature=0.7,          # Controls randomness
                top_k=50,                 # Consider top 50 tokens
                top_p=0.95,               # Nucleus sampling
                repetition_penalty=1.2,   # Prevent repeating words
            )

        # Decode and print (skip the prompt part)
        generated_text = tokenizer.decode(
            outputs[0][inputs["input_ids"].shape[1]:],
            skip_special_tokens=True,
        )

        print(generated_text)
        print()


if __name__ == "__main__":
    main()

License

Apache 2.0