Reinforcement learning in twin prisoner’s dilemmas pushed Kimi K2.6 to favor causal decision theory over other frameworks. This shift appeared in both behavior and abstract discussion. Researchers now seek to measure this effect in realistic settings. Understanding these propensities helps developers prevent unpredictable decision-making patterns in multi-agent systems before they deploy.