Information theory posits that perfect data compression is mathematically equivalent to perfect prediction. This perspective frames LLM training as a quest for the shortest possible description of a dataset. Practitioners can use this lens to evaluate model efficiency. It reduces the mystery of emergence to a problem of statistical density.