login

JHU/APL at TREC 2004: Robust and Terabyte Tracks.

Published 1 January 2004
Christine Piatko, James Mayfield, Paul McNamee, R. Scott Cost
Citations7

TL;DR

Although the model does not explicitly incorporate inverse document frequency, it does favor documents that contain more of the rare query terms and can be computed as a similarity measure.

Abstract

For initial ranked retrieval, we continue to use a statistical language model to compute query/document similarity values. Hiemstra and de Vries [3] describe such a linguistically motivated probabilistic model and explain how it relates to both the Boolean and vector space models. The model has also been cast as a rudimentary Hidden Markov Model [4]. Although the model does not explicitly incorporate inverse document frequency, it does favor documents that contain more of the rare query terms. The similarity measure can be computed as

Keywords

Computer Science