New STHTD-MP Method Speeds Off-Policy Prediction | dailyai.report
23 stories from today
Research
92d ago
New STHTD-MP Method Speeds Off-Policy Prediction
The STHTD-MP method replaces standard feature covariance metrics with the symmetric part of the behavior-policy Bellman matrix. This shift optimizes the update geometry in primal-dual saddle-point formulations. Researchers found this approach stabilizes gradient temporal-difference learning.
The Signal
Practitioners can now achieve faster convergence in off-policy prediction using linear function approximation without tuning multiple learning rates.