PUMA
Istituto di Scienza e Tecnologie dell'Informazione     
Silvestri F. Sorting out the document identifier assignment problem. In: Advances in Information Retrieval . 29th European Conference on IR Research, ECIR 2007 (Rome, 2-5 April 2007). Proceedings, pp. 101 - 112. Giambattista Amati, Claudio Carpineto and Giovanni Romano (eds.). (Lecture Notes in Computer Science, vol. 4425). Springer Verlag, 2007.
 
 
Abstract
(English)
The compression of Inverted File indexes in Web Search Engines has received a lot of attention in these last years. Compressing the index not only reduces space occupancy but also improves the overall retrieval performance since it allows a better exploitation of the memory hierarchy. In this paper we are going to empirically show that in the case of collections of Web Documents we can enhance the performance of compression algorithms by simply assigning identifiers to documents according to the lexicographical ordering of the URLs. We will validate this assumption by comparing several assignment techniques and several compression algorithms on a quite large document collection composed by about six million documents. The results are very encouraging since we can improve the compression ratio up to 40% using an algorithm that takes about ninety seconds to finish using only 100 MB of main memory.
URL: http://www.springerlink.com/content/y0755644n8n48627/fulltext.pdf
DOI: 10.1007/978-3-540-71496-5
Subject Identifier assignment
Indexing technique
Information retrieval index compression
H.3 Information Storage and Retrieval
H.3.1 Content Analysis and Indexing. Indexing Methods


Icona documento 1) Download Document PDF


Icona documento Open access Icona documento Restricted Icona documento Private

 


Per ulteriori informazioni, contattare: Librarian http://puma.isti.cnr.it

Valid HTML 4.0 Transitional