Engineers on Lobste.rs are debating the efficiency of systolic arrays versus dataflow architectures for tensor operations. The discussion highlights a growing frustration with memory bottlenecks in current GPU designs. These technical critiques suggest that software-defined hardware will likely supersede rigid chip layouts to improve inference speeds for practitioners.