TF-IDF, or Term Frequency-Inverse Document Frequency, is a numerical statistic used to evaluate how relevant a word is to a document in a collection or corpus. It's a product of two components:
- Term Frequency (TF): This measures how often a specific term appears within a single document. A higher TF indicates the term is more prominent in that document.
- Inverse Document Frequency (IDF): This measures how common or rare a term is across the entire corpus. Terms that appear in many documents have a lower IDF, while rarer terms have a higher IDF.
By multiplying TF and IDF, TF-IDF assigns a higher weight to terms that are frequent in a specific document but infrequent across the corpus, effectively highlighting words that are uniquely important to that document.