login

Document length normalization

Information Processing & ManagementPublished 1 September 1996Open access
Amit Singhal, Gerard Salton, Mandar Mitra, Chris Buckley
Citations202
SJR quartileQ1
SJR score2.06
SNIP2.91
View PDF

TL;DR

A modified technique is presented that attempts to match the likelihood of retrieving a document of a certain length to thelihood of documents of that length being judged relevant, and it is shown that this technique yields significant improvements in retrieval effectiveness.

Abstract

In the TREC collection—a large full-text experimental text collection with widely varying document lengths—we observe that the likelihood of a document being judged relevant by a user increases with the document length. We show that a retrieval strategy, such as the vector-space cosine match, that retrieves documents of different lengths with roughly equal chances, will not optimally retrieve useful documents from such a collection. We present a modified technique—pivoted cosine normalization—that attempts to match the likelihood of retrieving documents of all lengths to the likelihood of their relevance, and show that this technique yields significant improvements in retrieval effectiveness.

Keywords

Computer Science