Document length normalization
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A modified technique is presented that attempts to match the likelihood of retrieving a document of a certain length to thelihood of documents of that length being judged relevant, and it is shown that this technique yields significant improvements in retrieval effectiveness.
Abstract
In the TREC collection—a large full-text experimental text collection with widely varying document lengths—we observe that the likelihood of a document being judged relevant by a user increases with the document length. We show that a retrieval strategy, such as the vector-space cosine match, that retrieves documents of different lengths with roughly equal chances, will not optimally retrieve useful documents from such a collection. We present a modified technique—pivoted cosine normalization—that attempts to match the likelihood of retrieving documents of all lengths to the likelihood of their relevance, and show that this technique yields significant improvements in retrieval effectiveness.
