๐”– Bobbio Scriptorium
โœฆ   LIBER   โœฆ

Domain classification of technical terms using the Web

โœ Scribed by Mitsuhiro Kida; Masatsugu Tonoike; Takehito Utsuro; Satoshi Sato


Publisher
John Wiley and Sons
Year
2007
Tongue
English
Weight
480 KB
Volume
38
Category
Article
ISSN
0882-1666

No coin nor oath required. For personal study only.

โœฆ Synopsis


Abstract

This paper proposes a method of domain classification of technical terms using the Web. In the proposed method, it is assumed that, for a certain technical domain, a list of known technical terms of the domain is given. Technical documents of the domain are collected through the Web search engine, which are then used for generating a vector space model for the domain. The domain specificity of a target term is estimated according to the distribution of the domain of the sample pages of the target term. Experimental evaluation results show that the proposed method of domain classification of a technical term achieved mostly 90% precision/recall. We then apply this technique of estimating domain specificity of a term to the task of discovering novel technical terms that are not included in any existing lexicons of technical terms of the domain. Out of 1000 randomly selected candidates of technical terms per domain, we discovered about 100 to 200 novel technical terms. ยฉ 2007 Wiley Periodicals, Inc. Syst Comp Jpn, 38(14): 11โ€“19, 2007; Published online in Wiley InterScience (www.interscience.wiley.com). DOI 10.1002/scj.20852


๐Ÿ“œ SIMILAR VOLUMES


Automatic classification of Web resource
โœ Charlotte Jenkins; Mike Jackson; Peter Burden; Jon Wallis ๐Ÿ“‚ Article ๐Ÿ“… 1998 ๐Ÿ› Elsevier Science ๐ŸŒ English โš– 403 KB

The Wolverhampton Web Library' (WWLib) is a World Wide Web search engine that provides access to UK based intinmation. The experimental version, developed in IYY5. was a success but highlighted the need for a much higher degree of automation. An interesting feature of the experimental WWLib was that

Information Science in the web era: A te
โœ Fidelia Ibekwe-SanJuan ๐Ÿ“‚ Article ๐Ÿ“… 2009 ๐Ÿ› Wiley (John Wiley & Sons) ๐ŸŒ English โš– 190 KB

## Abstract We propose a methodology for mapping the research in Information Science (IS) field based on a combined use of symbolic (linguistic) and numeric information. Using the same list of 12 IS journals as in earlier studies on this same topic (White & McCain 1998; Zhao & Strotmann 2008a&b), w

On the Classification of Stable Domains
โœ Bruce Olberding ๐Ÿ“‚ Article ๐Ÿ“… 2001 ๐Ÿ› Elsevier Science ๐ŸŒ English โš– 151 KB

In the second half of a two-part study of stable domains, we explore the extent to which the stability of an integral domain is determined by the stability of its prime and finitely generated ideals. This yields pullback theorems for stable domains, a method of constructing nonstandard examples of s

How is science cited on the Web? A class
โœ Kayvan Kousha; Mike Thelwall ๐Ÿ“‚ Article ๐Ÿ“… 2007 ๐Ÿ› John Wiley and Sons ๐ŸŒ English โš– 279 KB ๐Ÿ‘ 1 views

## Abstract Although the analysis of citations in the scholarly literature is now an established and relatively well understood part of information science, not enough is known about citations that can be found on the Web. In particular, are there new Web types, and if so, are these trivial or pote