ai-sherpa/clipped-q-learning-robust-repro
rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound
rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound
rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound
rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound
rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound
rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound
rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound
Update logbook: Repro - Clipped Q-Learning: Your Value Clipping Is Secretly A Robust Operator
initial commit
