STARFlow2 integrates autoregressive normalizing flows directly into Transformer architectures to generate interleaved text and images. This approach avoids the visual fidelity loss common in discrete tokenization. By sharing the same causal mask and KV-cache as LLMs, it eliminates the structural asymmetry of diffusion-based systems. Practitioners gain a unified framework for high-fidelity multimodal output.