makas8585/dual-brain-inference
Design of Inference Engine Architecture via Shared KV-Cache Based Real-Time Dual-Attention Ensemble makas8585Independent Researchertjdrldud850@gmail.com Abstract Transformer-based large language models (LLMs) are fundamentally vulnerable to self-bias and hallucination due to their reliance on a single autoregressive attention mechanism. The same component that generates tokens also implicitly validates them, creating a closed loop that cannot self-correct… See the full description on the dataset page: https://huggingface.co/datasets/makas8585/dual-brain-inference.
016
