Late-interaction retrieval encodes queries and documents as token vectors. At search time, it takes the best document-token match for each query token and sums those maxima.
This preserves fine-grained clues that a single document vector can lose. Storing token vectors and comparing many pairs increases index size and compute cost.
When to use
Use it when each detail in a query should contribute to document relevance.