STHTD-MP Speeds Up Off-Policy Prediction | dailyai.report
23 stories from today
Research
92d ago
STHTD-MP Speeds Up Off-Policy Prediction
The STHTD-MP method replaces standard feature covariance metrics with the symmetric part of the behavior-policy Bellman matrix. This change optimizes the update geometry in primal-dual saddle-point formulations. By utilizing behavior-policy transition information, the approach accelerates convergence in linear function approximation.
The Signal
It offers a more efficient alternative for stable off-policy prediction.