forkjoin-ai/buleyean-qwen2.5-0.5b
036
buleyean-qwen2.5-0.5b
Buleyean RL -- trained on what is NOT rather than positive reinforcement.
No reward model. No chosen examples. The complement distribution derived from rejection counts alone is the training target.
Model Details
What is Buleyean RL?
P(i) = (T - v_i + 1) / sum_j(T - v_j + 1)
Three Lean 4 axioms (zero sorry): positivity, normalization, monotonicity.
Loss: L = 0.7 * KL(P_bule || P_model) + 0.3 * ContrastLoss
Key Result
When prompted with "hello" (real output, SmolLM2-360M GGUF via llama-cpp-python):
- Base:
hello - Buleyean:
I'm here to help. What's on your mind?
Whitepaper
[Proof of Life: Bottling Infinity in Distributed Systems -- φ² = φ + 1](https://forkracefold.com/)
500+ Lean 4 theorems. Zero sorry markers. Section 15.29 covers Buleyean RL. Chapter 29 is the full treatment.
Links
- Library | Demo | Data
- Whitepaper | MPL-2.0
