Engineers on Lobste.rs are debating the efficiency of systolic arrays versus dataflow architectures for LLM inference. The discussion focuses on memory bandwidth bottlenecks and the energy cost of data movement. These technical trade-offs determine whether future ASICs can reduce power consumption. Practitioners should track these architectural shifts to optimize hardware-aware model quantization.