XMill
ACM SIGMOD RecordPublished 16 May 2000Open access
Hartmut Liefke, Dan Suciu
Citations418
SJR quartileQ2
SJR score0.69
SNIP0.92
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
We describe a tool for compressing XML data, with applications in data exchange and archiving, which usually achieves about twice the compression ratio of gzip at roughly the same speed. The compressor, called XMill, incorporates and combines existing compressors in order to apply them to heterogeneous XML data: it uses zlib, the library function for gzip, a collection of datatype specific compressors for simple data types, and, possibly, user defined compressors for application specific data types.
Keywords
Computer Science
Mathematics and Computers in SimulationIntroduction to automata theory, languages and computation
10,827 Citations1981
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
IEEE Transactions on Information TheoryA universal algorithm for sequential data compression
5,467 Citations1977J. Ziv, A. Lempel
The compression ratio achieved by the proposed universal code uniformly approaches the lower bounds on the compression ratios attainable by block-to-variable codes and variable- to-block codes designed to match a completely specified source.
A Block-sorting Lossless Data Compression Algorithm
2,362 Citations1994Michael T. Burrows, D. J. Wheeler
A block-sorting, lossless data compression algorithm, and the implementation of that algorithm, that achieves speed comparable to algorithms based on the techniques of Lempel and Ziv, but obtains compression close to the best statistical modelling techniques.
KybernetesData Compression: The Complete Reference
1,599 Citations2009
Detailed descriptions and explanations of the most well-known and frequently used compression methods are covered in a self-contained fashion, with an accessible style and technical level for specialists and nonspecialists.
DataGuides: Enabling Query Formulation and Optimization in Semistructured Databases
1,147 Citations1997Roy Goldman, Jennifer Widom
The theoretical foundations of DataGuides are presented along with an algorithm for their creation and an overview of incremental maintenance, and performance results based on the implementation of dataGuides in the Lore DBMS for semistructured data are provided.
Nucleic Acids ResearchThe EMBL data library
312 Citations1988Graham Cameron
The contents of the database, how it is available, and possible future enhancements of Data Library services are described.
Information Processing & ManagementA new challenge for compression algorithms: Genetic sequences
273 Citations1994Stéphane Grumbach, Fariza Tahi
A lossless algorithm is presented, biocompress-2, to compress the information contained in DNA and RNA sequences, based on the detection of regularities, such as the presence of palindromes, which leads to the highest compression of DNA.
Data compression
249 Citations1997David Salomon
Compressing relations and indexes
207 Citations2002J. Goldstein, R. Ramakrishnan +1 more
A new compression algorithm that is tailored to database applications that can be applied to a collection of records, and is especially effective for records with many low to medium cardinality fields and numeric fields, is proposed.
ACM SIGMOD RecordDatabase compression
138 Citations1993Mark A. Roth, Scott J. Van Horn
This work addresses several aspects of reversible data compression and compression techniques, including general concepts of data compression; a number of compression techniques; a comparison of the effects of compression on common data types; advantages and disadvantages; and future research needs.
ACM SIGMOD RecordInferring structure in semistructured data
119 Citations1997Svetlozer Nestorov, Serge Abiteboul +1 more
A notion of a type hierarchy for such data is proposed, and a method for deriving the type hierarchy is outlined, and rules for assigning types to data elements are outlined.
Data Compression Support in Databases
94 Citations1994Balakrishna R. Iyer, David Wilhite
Various design issues arise in the use of data compression in the dbms from the choice of algorithm, statistics collection, hardware versus software based compression, location of the compression function in the overall computer system architecture, unit of compression, update in place, and the application of log’ to compressed data.
IEEE Transactions on Knowledge and Data EngineeringBlock-oriented compression techniques for large statistical databases
71 Citations1997Wee Keong Ng, Chinya V. Ravishankar
The authors explore the compression of large statistical databases and propose techniques for organizing the compressed data such that standard database operations such as retrievals, inserts, deletes and modifications are supported.
Nucleic Acids ResearchThe EMBL Data Library
40 Citations1992Desmond G. Higgins, Rainer Fuchs +2 more
The principal role of the EMBL Data Library has been to maintain and distribute a database of nucleotide sequences (the EMBL Nucleotide Sequence Database), which also supports and maintains the protein sequence database SWISS-PROT and distributes other databases of interest to molecular biologists.
ACM SIGMOD RecordAn extensible compressor for XML data
18 Citations2000Hartmut Liefke, Dan Suciu
This paper describes XMilI's extensible architecture and summarizes some other aspects of XMill, such as container expressions and an experimental evaluation, to be used in data exchange and archiving.
