wang2226/steering-showcase
0
Send depth, token limit and contrast continuations; keep live takeaways neutral
Reject the last decoder block as a steering layer
Report the GPU it ran on, and take the softmax on the CPU in float64
Record the CPU/GPU comparison result in the module docstrings
Read probabilities from the full softmax, and report the revision
Serve several models, and a visitor's own prompt
Cache deterministic results so a visitor costs no GPU quota
Deploy steering showcase API
initial commit
