We study whether transformers can learn to implicitly reason over parametric knowledge, a skill that even the most capable language models struggle with.
We derive the distortion-rate function for this setup as a linear program, and provide an efficient algorithm to compute this fundamental limit via the dual of the linear program.