SupraLabs/Supra-Router-51M
1482.3k
1---2library_name: transformers3tags:4- router5- orchestrator6- slm7- edge-computing8- mixture-of-experts9- text-generation10pipeline_tag: text-generation11model_type: llama12datasets:13- SupraLabs/Prompt-Routing-Dataset14language:15- en16base_model:17- SupraLabs/Supra-1.5-50M-Base-exp18---19 20<h1 align="center">Supra-Router-51M · Multi-Task Infrastructure Routing Model</h1>21 2223 24<h2 align="center">About the Model</h2>25 26GGUF model [here](https://huggingface.co/SupraLabs/Supra-Router-51M-gguf)27 28**Supra-Router-51M** is an ultra-lightweight, high-speed infrastructure traffic controller optimized for localized edge orchestration. With only **51.7 million parameters**, this micro-LLM acts as a defensive gateway for multi-model ecosystems, accurately determining when user requests can be processed locally by an Edge SLM or when they must be triaged to a cloud-hosted frontier intelligence layer.29 30The model was built by fine-tuning a pre-trained 51M base on the `SupraLabs/Prompt-Routing-Dataset` (992 rows). Rather than acting as a naive binary classifier, the model uses **Multi-Task Sequence Generation** to map out the underlying properties of a prompt before predicting the final routing token, anchoring its attention heads to robust language and structural logic features.31 32---33 34## Multi-Task Decision Sequence35 36To run inference, wrap your user query inside the structural framing tokens used during training (`Task: [Prompt]\nAnalysis: `). The model will output a deterministic, pipe-separated string containing the full telemetry of the prompt's cognitive requirements:37 38### Expected Output Target Schema:39```text40Domain: [Semantic Field] | Complexity: [1-5] | Math: [True/False] | Code: [True/False] | Route: [small model/big model] | Justification: [Rule-driven infrastructure reasoning]41```42 43## Why this works:44 45By forcing a sub-100M parameter model to calculate the semantic domain, structural complexity, and technical flags before it emits the final Route token, the network effectively runs an internal feature-activation map. This multi-task sequence prevents localized weight collapse and guarantees stable routing boundaries.46 47## Training Telemetry & Optimization48 49- Dataset Source: SupraLabs/Prompt-Routing-Dataset (992 samples)50- Training Duration: 5 Epochs51- Checkpoint Selection: Peak generalization was reached during Epoch 3 (eval_loss: 0.1342). To eliminate late-stage micro-model memorization and validation drift, the training state was automatically rewound and saved at this numerical peak.52- Precision: bfloat1653- Hardware Footprint: Optimized sequence processing length of 3840 tokens, ensuring rapid inference execution with negligible CPU/GPU overhead (sub-millisecond generation speeds).54 55## Inference & Gateway Implementation56 57Use this direct script to test or wrap the model inside a live production orchestrator or FastAPI gateway. It enforces greedy decoding (do_sample=False) for maximum decision stability.58 59```python60import torch61from transformers import AutoModelForCausalLM, AutoTokenizer62 63MODEL_ID = "SupraLabs/Supra-Router-51M"64 65print("[*] Initializing local infrastructure router...")66tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)67model = AutoModelForCausalLM.from_pretrained(68 MODEL_ID,69 dtype=torch.bfloat16,70 device_map="auto"71)72model.eval()73 74# Example prompt showcasing keyword-trap evasion75user_prompt = "Write a movie script about a chef who gets lost at sea."76 77# Format to match internal SFT attention alignment78formatted_input = f"Task: {user_prompt}\nAnalysis: "79inputs = tokenizer(formatted_input, return_tensors="pt").to(model.device)80 81with torch.no_grad():82 outputs = model.generate(83 **inputs,84 max_new_tokens=128,85 do_sample=False, 86 pad_token_id=tokenizer.pad_token_id,87 eos_token_id=tokenizer.eos_token_id88 )89 90generated_ids = outputs[0][inputs["input_ids"].shape[1]:]91print(tokenizer.decode(generated_ids, skip_special_tokens=True).strip())92```93 94## Proven Benchmarks & Defensive Boundaries95 96During edge validation testing, Supra-Router-51M demonstrated robust resilience against adversarial prompt strings:97 98- Keyword Trap Evasion: Successfully identifies semantic context rather than matching tokens. Prompts containing words like "script" or "calculus" are correctly parsed as creative writing (not programming/math code) and routed locally to the small model when complexity is low.99- Complexity-Driven Safety Net: In instances where programming syntax or technical boundaries are ambiguous (e.g., complex regex or architectural database frames), the model naturally scales its evaluation metrics to Complexity: 3, automatically triggering a big model route override.100- Deterministic Offloading: Safely captures multi-step logic paths, calculus concepts, and code generation scripts, instantly assigning them to cloud-scale frontier endpoints.