✦ LIBER ✦

High-speed rough clustering for very large document collections

✍ Scribed by Kazuaki Kishida

Publisher: John Wiley and Sons
Year: 2010
Tongue: English
Weight: 249 KB
Volume: 61
Category: Article
ISSN: 1532-2882
DOI: 10.1002/asi.21311

No coin nor oath required. For personal study only.

✦ Synopsis

Abstract

Document clustering is an important tool, but it is not yet widely used in practice probably because of its high computational complexity. This article explores techniques of high‐speed rough clustering of documents, assuming that it is sometimes necessary to obtain a clustering result in a shorter time, although the result is just an approximate outline of document clusters. A promising approach for such clustering is to reduce the number of documents to be checked for generating cluster vectors in the leader–follower clustering algorithm. Based on this idea, the present article proposes a modified Crouch algorithm and incomplete single‐pass leader–follower algorithm. Also, a two‐stage grouping technique, in which the first stage attempts to decrease the number of documents to be processed in the second stage by applying a quick merging technique, is developed. An experiment using a part of the Reuters corpus RCV1 showed empirically that both the modified Crouch and the incomplete single‐pass leader–follower algorithms achieve clustering results more efficiently than the original methods, and also improved the effectiveness of clustering results. On the other hand, the two‐stage grouping technique did not reduce the processing time in this experiment.

📜 SIMILAR VOLUMES

Topic modeling for mediated access to ve

Topic modeling for mediated access to very large document collections

✍ Gheorghe Muresan; David J. Harper 📂 Article 📅 2004 🏛 John Wiley and Sons 🌐 English ⚖ 457 KB

Incremental clustering for very large do

Incremental clustering for very large document databases: Initial MARIAN Experience

✍ Fazli Can; Edward A. Fox; Cory D. Snavely; Robert K. France 📂 Article 📅 1995 🏛 Elsevier Science 🌐 English ⚖ 727 KB

Non-cryopreserved peripheral blood proge

Non-cryopreserved peripheral blood progenitor cells collected by a single very large-volume leukapheresis: A simplified and effective procedure for support of high-dose chemotherapy

✍ Christos A. Papadimitriou; Meletios A. Dimopoulos; Vassilios Kouvelis; Evangelos 📂 Article 📅 2000 🏛 John Wiley and Sons 🌐 English ⚖ 38 KB 👁 2 views

High-dose chemotherapy with autologous peripheral blood progenitor cell (PBPC) support has become a widely used treatment strategy. In order to simplify the procedure, a single very large-volume leukapheresis programme combined with short-term refrigerated storage of the PBPC was developed. Seventy-