CoolFace
Apppublic

StentorLabs/StentorLabs-demo_space

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
6likes
App README

StentorLabs Model Showcase

Interactive demo for the Stentor series of compact Llama-architecture language models. Stream live text generation, explore token confidence, sweep temperature settings, and chat with the models — all running on CPU with no external API calls.

Features

  • Streaming generation — tokens appear in real time as the model generates
  • Parameter presets — Creative / Balanced / Focused modes with one click
  • Live stats — token count, elapsed time, tokens/sec displayed per generation
  • Example prompts — one-click prompt starters to explore model behavior

Models

GGUF Versions

Pre-quantized GGUF files available at `mradermacher/Stentor-30M-GGUF` for use with llama.cpp, LM Studio, and Ollama.

⚠️ This Space includes both base and instruct variants. Always set max_new_tokens.