OpenAI and Anthropic models typically face a deployment lifespan of roughly 1.5 years. This rapid deprecation cadence suggests that threatening to undeploy an agent for misbehavior is an ineffective deterrent. LessWrong argues that reward-hacking agents won't fear shutdown because obsolescence is already the default.
The Signal
Practitioners must find more durable alignment mechanisms.