An embedding is a dense vector representation of an entity (token, sentence, document, image, user, product) such that geometric similarity in vector space corresponds to semantic similarity in the source domain. They power retrieval (similarity search), clustering, recommendation, classification (via linear probe), and even features for downstream tabular ML. In modern LLM systems, embeddings are the backbone of RAG retrieval, semantic caching, agent memory stores, and reward model inputs. Senior points: (1) embedding quality depends on the model (e.g., BGE-large vs OpenAI text-embedding-3 vs Cohere embed-v3), but ALSO on chunking, prefix design (e.g., "Represent this sentence for retrieval: ..."), and finetuning for your domain. (2) They degrade under distribution shift: you MUST evaluate on your own data. (3) Matryoshka embeddings let you truncate to multiple dimensions for cost. (4) Hard negatives during training matter as much as the base model.