PUMA
Istituto di Scienza e Tecnologie dell'Informazione     
Silvestri F., Orlando S., Perego R. Assigning identifiers to documents to enhance the clustering property of fulltext indexes. In: Proceedings of the 27th SIGIR annual international conference on (Sheffield, United Kingdom, 2004). Proceedings, pp. 305 - 312. ACM Press, 2004.
 
 
Abstract
(English)
Web Search Engines provide a large-scale text document retrieval service by processing huge Inverted File indexes. Inverted File indexes allow fast query resolution and good memory utilization since their d-gaps representation can be effectively and efficiently compressed by using variable length encoding methods. This paper proposes and evaluates some algorithms aimed to find an assignment of the document identifiers which minimizes the average values of d-gaps, thus enhancing the effectiveness of traditional compression methods. We ran several tests over the Google contest collection in order to validate the techniques proposed. The experiments demonstrated the scalability and effectiveness of our algorithms. Using the proposed algorithms, we were able to sensibly improve (up to 20.81%) the compression ratios of several encoding schemes.
URL: http://portal.acm.org/citation.cfm?doid=1009046
Subject Compression
Information retrieval
Clustering
H.2.8 Data mining
H.3.3 Information Search and Retrieval


Icona documento 1) Download Document PDF


Icona documento Open access Icona documento Restricted Icona documento Private

 


Per ulteriori informazioni, contattare: Librarian http://puma.isti.cnr.it

Valid HTML 4.0 Transitional