CoolFace
Apppublic

ai-sherpa/clipped-q-learning-robust-repro

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
9 commits on main
b4357462mo ago

rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound

ai-sherpa
2b888f62mo ago

rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound

ai-sherpa
ded8a832mo ago

rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound

ai-sherpa
ea40ae62mo ago

rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound

ai-sherpa
c86591b2mo ago

rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound

ai-sherpa
a00dd702mo ago

rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound

ai-sherpa
10fba202mo ago

rework toy->verified: worst-case-LP robust operator, contraction+fixed point, N=5 unbiasedness, order-one-constant sub-linear regret under theorem bound

ai-sherpa
59a1df52mo ago

Update logbook: Repro - Clipped Q-Learning: Your Value Clipping Is Secretly A Robust Operator

ai-sherpa
428b1ca2mo ago

initial commit

ai-sherpa