New Mirror-Prox Method Speeds Off-Policy Prediction | dailyai.report
23 stories from today
Research
92d ago
New Mirror-Prox Method Speeds Off-Policy Prediction
The STHTD-MP method replaces standard covariance metrics with the symmetric part of the behavior-policy Bellman matrix. This change optimizes the update geometry in primal-dual saddle-point formulations. It simplifies tuning by using a single learning rate for both primal and auxiliary variables.
The Signal
Researchers can now achieve faster convergence in stable off-policy prediction with linear function approximation.