Liu et al. (2023) "Lost in the Middle" showed that even at long context lengths, LLMs reliably use information at the very beginning and end of the context but struggle to recall information in the middle: a U-shaped recall curve. This is partly because long contexts dilute attention scores (each retrieved chunk gets a smaller share of softmax mass) and partly because training data is biased toward short contexts. Mitigation strategies: (1) retrieval re-ranking to put most relevant chunks at the ends; (2) step-back prompting or chain-of-thought that re-attends to the middle; (3) prepending summaries; (4) compressing the "middle" via summarization; (5) using models trained with attention sinks or specialized long-context training (e.g., YaRN). For RAG, this is why simple top-k retrieval often under-performs hybrid + reranker pipelines. Senior caveat: the phenomenon is partially mitigated in newer models trained with long-context targets, but not eliminated.