dylantom2012/fly-head-potion-8m
Fly head — a text classifier whose decision layer is eight floats
{"proto": 3.7233, "desc": 0.4448, "knn1": 3.0027, "knnk": 4.1251,
"margin": 2.4977, "prior": 0.4307, "white": 1.5815, "tight": 0.6906}That is the entire trained decision layer. It sits on top of `minishlab/potion-base-8M`, a static embedding — no attention, no matrix multiply in the forward pass, just a lookup table and an average.
Measured on 10,000 held-out examples across four datasets:
For reference, typesafe/jev — a hosted commercial decision model — scores 78.8 % on the same 10,000 items at 381 ms per call. Jev needs no labels; this head needs ~2,000 per task. That difference is the whole story, and it is discussed honestly in the benchmark repository.
Per-task accuracy
Use it
pip install numpy model2vec
python inference.py banking77 "I still haven't received the card I ordered"
# 45.3% card arrival
# 7.6% pending card payment
# 6.8% card delivery estimatefrom inference import FlyHead
head = FlyHead("banking77")
head("my contactless stopped working") # -> [(option, probability), ...]How it works
The head scores every (query, option) pair with eight normalised features and takes the argmax — the same shape as a ranking function, which is why it supports any option set rather than a fixed output layer:
The eight weights were found with a (1+λ) evolution strategy, λ=14 with σ-restarts, maximising 0.7 × mean accuracy + 0.3 × worst-task accuracy across the four tasks. Weights are selected on the training objective, never on validation score.
heads/<dataset>.npz holds the offline state the features need: class centroids, 1,500 fit vectors for the kNN features, option-description vectors, the global mean and the class log-priors. About 1.4 MB per task.
Limitations
- It needs labelled examples (~2,000 per task). For a genuinely cold start, use a zero-shot cross-encoder instead — the benchmark repo measures that path too (78.7 % macro, 50–616 ms, CPU).
- Adding a new option requires recomputing that class's centroid, which needs examples of it. Nothing needs retraining, but you do need data.
- Single seed, single machine. No confidence intervals. See the repo's limitations.
- Trained and evaluated on four English datasets. Nothing here has been tested in another language.
Citation / provenance
Built and measured in `zhlei07/open-system-one`. The direction came from the repository owner; every experiment was designed and run by Claude Code (Opus 5) in a single session, including the ideas that failed. All raw outputs are published so the numbers can be checked rather than trusted.
Code: MIT. The base encoder and the datasets retain their own licences.
