Metrics for the scope of a collection
β Scribed by Robert B. Allen; Yejun Wu
- Publisher
- John Wiley and Sons
- Year
- 2005
- Tongue
- English
- Weight
- 79 KB
- Volume
- 56
- Category
- Article
- ISSN
- 1532-2882
No coin nor oath required. For personal study only.
β¦ Synopsis
Abstract
Some collections cover many topics, while others are narrowly focused on a limited number of topics. We introduce the concept of the βscopeβ of a collection of documents and we compare two ways of measuring it. These measures are based on the distances between documents. The first uses the overlap of words between pairs of documents. The second measure uses a novel method that calculates the semantic relatedness to pairs of words from the documents. Those values are combined to obtain an overall distance between the documents. The main validation for the measures compared Web pages categorized by Yahoo. Sets of pages sampled from broad categories were determined to have a higher scope than sets derived from subcategories. The measure was significant and confirmed the expected difference in scope. Finally, we discuss other measures related to scope.
π SIMILAR VOLUMES
Julie Hyzy, the New York Times bestselling author of the White House Chef Mysteries and the Manor House Mystery series, brings together nine of her most popular short stories, re-released in this new collection. From subtle to edgy, from gentle to terrifying, these fast-paced stories explore the spe
We thank Rui Sousa for the gift of the Y639F T7 RNA polymerase plasmid, Kara Juneau and Tom Cech for the DC209 P4-P6 plasmid, Cecilia Cortez and Shirshendu Deb for assistance with nucleoside synthesis, Joe Olvera for technical assistance, and Cheryl Small for assistance with manuscript preparation.