Compression: a key for next-generation text retrieval systems
ComputerPublished 1 January 2000Open access
Nívio Ziviani, Edleno Silva de Moura, Gonzalo Navarro, Ricardo Baeza‐Yates
Citations138
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The authors' technique combines several data compression features to provide economical storage, faster indexing, and accelerated searches.
Abstract
The continually growing Web challenges information retrieval systems to deliver data quickly. The authors' technique combines several data compression features to provide economical storage, faster indexing, and accelerated searches.
Keywords
Computer Science
Modern Information Retrieval
11,544 Citations1999Ricardo Baeza‐Yates, Berthier Ribeiro‐Neto
Computer Architecture: A Quantitative Approach
9,536 Citations1989John L. Hennessy, David A. Patterson
This best-selling title, considered for over a decade to be essential reading for every serious student and practitioner of computer design, has been updated throughout to address the most important trends facing computer designers today.
Proceedings of the IREA Method for the Construction of Minimum-Redundancy Codes
6,200 Citations1952David A. Huffman
NatureAccessibility of information on the web
1,356 Citations1999Steve Lawrence, C. Lee Giles
As the web becomes a major communications medium, the data on it must be made more accessible, and search engines need to make the data more accessible.
Communications of the ACMFast text searching
728 Citations1992Sun Wu, Udi Manber
T h e string-matching problem is a very c o m m o n problem; there are many extensions to t h i s problem; for example, it may be looking for a set of patterns, a pattern w i t h "wi ld cards," or a regular expression.
Communications of the ACMA new approach to text searching
598 Citations1992Ricardo Baeza‐Yates, Gastón H. Gonnet
A family of simple and fast algorithms for solving the classical string matching problem, string matching with don't care symbols and complement symbols, and multiple patterns are introduced.
Information Processing & ManagementOverview of the Second Text Retrieval Conference (TREC-2)
310 Citations1995Donna Harman
This conference, co-sponsored by ARPA and NIST, brought together information retrieval researchers to discuss their system results on the new TIPSTER test collection, and represented a breakthrough in cross-system evaluation in information retrieval.
ACM Transactions on Information SystemsFast and flexible word searching on compressed text
246 Citations2000Edleno Silva de Moura, Gonzalo Navarro +2 more
A fast compression technique for natural language texts that allows a large number of variations over the basic word and phrase search capability, such as sets of characters, arbitrary regular expressions, and approximate matching.
AlgorithmicaFaster Approximate String Matching
167 Citations1999Ricardo Baeza‐Yates
A new algorithm for on-line approximate string matching based on the simulation of a nondeterministic finite automaton built from the pattern and using the text as input, which is among the fastest for typical text searching, being the fastest in some cases.
Information RetrievalAdding Compression to Block Addressing Inverted Indexes
103 Citations2000Gonzalo Navarro, Edleno Silva de Moura +3 more
This work presents a compressed inverted file that indexes compressed text and uses block addressing, and compares the index against three separate techniques for varying block sizes, showing that the index is superior to each isolated approach.
Document Filtering for Fast Ranking
28 Citations1994Michael Persin
The experiments show that the proposed evaluation technique reduces both main memory usage and query evaluation time, based on early recognition of which documents are likely to be highly ranked, without degradation in retrieval effectiveness.
