The "expenditure horizon" metric calculates the exact point where METR agents become more expensive than human workers. Early tests on the NanoGPT speedrun show underwhelming results. The framework currently suffers from significant blind spots. Practitioners should view this as an incremental attempt to quantify agent efficiency rather than a definitive industry standard.