○ unseen · kind algorithm · level 0 · 0h
- Aplica: Natural Language Processing
- Contradice: Embeddings
Weights each term in a document by how often it appears here (TF) times how rare it is across the whole corpus (IDF) — a classic pre-embeddings way to turn text into numbers.
Mecanismo. tf-idf(t,d) = tf(t,d) × log(N / df(t)) — df(t) is the number of documents containing term t, N the corpus size. Common words (“the”) appear in nearly every document → near-zero IDF → near-zero weight regardless of how often they repeat. Rare-but-repeated words get high weight. Produces a sparse vector per document usable for search ranking or document similarity (cosine). Still the backbone of BM25 — the ranking function behind most keyword search engines (Elasticsearch included) — which adds document-length normalization and term-frequency saturation on top of the same idea.
Ejercicio. h_comment_mining.R’s wmjq_result_29/term_freq.txt is the raw term-frequency step — the direct precursor to TF-IDF in the same pipeline (it stops at TF; adding the IDF weighting is the natural next exercise).
Enlaces
- Aplica: Natural Language Processing
- Contradice: Embeddings