policy-learning
policy_pos_neg_2012_manifesto_batch32_50epoch_learningdefaultpolicy__2012_manifesto_batch64_100epoch_learningdefaultpolicy__2612_manifesto_batch32_20epoch_learningdefaultpolicy__2012_manifesto_batch32_50epoch_learningdefaulteval2_final_policyeval1_final_policytaxi-Q-learning-off-policytaxi-Q-learning-on-policy
mw_policy_learningrepro-off-policy-learning-in-large-action-spaces-optimization-matters-more-than-estimation
Reproduction: Off-Policy Learning in Large Action Spaces - Optimization Matters More Than Estimation
Paper Information
Title: Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
OpenReview ID: srIStBTJiu
Conference: ICML 2026
Task: Compare optimization landscapes of IPS vs PWLL for off-policy policy learning
Reproduction Summary
This reproduction evaluates the paper's core thesis: optimization landscape (not… See the full description on the dataset page: https://huggingface.co/datasets/sabaridsnfuji/repro-off-policy-learning-in-large-action-spaces-optimization-matters-more-than-estimation.
From-911-Calls-to-Policy-Gradients-Reinforcement-Learning-for-Public-Safety-Dispatchoff-policy-learning-in-large-action-spaces-reprorepro-learning-human-robot-collaboration-via-heterogeneous-agent-lyapunov-policy-optimizatiorepro-milestone-guided-policy-learning-for-long-horizon-language-agentsrepro-zero-shot-off-policy-learningrepro-speedup-patch-learning-a-plug-and-play-policy-to-accelerate-embodied-manipulation
