barha/granite-switch-4.0-350m-demo
Granite Switch 4.0 350M — 3-Adapter Demo
A single Granite Switch checkpoint built on `ibm-granite/granite-4.0-350m` with three task LoRA adapters embedded in one model. Each adapter is activated by a control token, so one deployed checkpoint serves three different tasks — no adapter swapping, no separate model loads.
This is the multi-adapter companion to the single-adapter `barha/granite-switch-4.0-350m-cti`.
Embedded adapters
All three adapters share an identical LoRA shape (rank 16, alpha 32, on the fused q/k/v/o attention projections and the input_linear / output_linear MLP projections), which is what lets them stack cleanly into one switch checkpoint.
Control token IDs: 100352, 100353, 100354 (3 new tokens; vocab 100355).
Per-adapter evaluation (on granite-4.0-350m)
These are the standalone scores of each source LoRA; embedding them in the switch does not change adapter weights.
How adapter selection works
Granite Switch routes per request via the control token: place the adapter's control token in the prompt (the chat template handles placement) and the switch activates that adapter for the turn. With no control token, the base model runs unmodified.
Usage
Compose / inference follow the standard Granite Switch flow — see the granite-switch repo and its tutorials. The checkpoint loads as a GraniteSwitchForCausalLM; adapter_index.json lists the adapter → control-token mapping and io_configs/<adapter>/io.yaml carries each adapter's I/O contract.
Build
Composed with granite_switch.composer.compose_granite_switch:
python -m granite_switch.composer.compose_granite_switch \
--base-model ibm-granite/granite-4.0-350m \
--technology lora \
--adapters cti-technique-mapping text-to-json genai-attack-vector \
--output ./granite-switch-4.0-350m-demoBase params 352M → composed 359M (+2.0%). See BUILD.md and compose_report.json in this repo for the full composition report.
License: Apache-2.0
