Istituto di Scienza e Tecnologie dell'Informazione     
Silvestri F., Orlando S., Perego R. Assigning identifiers to documents to enhance the clustering property of fulltext indexes. In: Proceedings of the 27th SIGIR annual international conference on (Sheffield, United Kingdom, 2004). Proceedings, pp. 305 - 312. ACM Press, 2004.
Web Search Engines provide a large-scale text document retrieval service by processing huge Inverted File indexes. Inverted File indexes allow fast query resolution and good memory utilization since their d-gaps representation can be effectively and efficiently compressed by using variable length encoding methods. This paper proposes and evaluates some algorithms aimed to find an assignment of the document identifiers which minimizes the average values of d-gaps, thus enhancing the effectiveness of traditional compression methods. We ran several tests over the Google contest collection in order to validate the techniques proposed. The experiments demonstrated the scalability and effectiveness of our algorithms. Using the proposed algorithms, we were able to sensibly improve (up to 20.81%) the compression ratios of several encoding schemes.
URL: http://portal.acm.org/citation.cfm?doid=1009046
Subject Compression
Information retrieval
H.2.8 Data mining
H.3.3 Information Search and Retrieval

Icona documento 1) Download Document PDF

Icona documento Open access Icona documento Restricted Icona documento Private


Per ulteriori informazioni, contattare: Librarian http://puma.isti.cnr.it

Valid HTML 4.0 Transitional