A new replication study on arXiv challenges the use of Floating Point Operations (FLOPs) as a proxy for computational cost. Researchers found that layers with identical FLOP counts execute at different speeds due to varying parallelization efficiency. This discrepancy makes traditional efficiency formulas unreliable on modern hardware. Practitioners should prioritize actual wall-clock time over theoretical FLOPs.