๐”– Scriptorium
โœฆ   LIBER   โœฆ

๐Ÿ“

Explorations in Automatic Thesaurus Discovery

โœ Scribed by Gregory Grefenstette (auth.)


Publisher
Springer US
Year
1994
Tongue
English
Leaves
312
Series
The Springer International Series in Engineering and Computer Science 278
Edition
1
Category
Library

โฌ‡  Acquire This Volume

No coin nor oath required. For personal study only.

โœฆ Synopsis


Explorations in Automatic Thesaurus Discovery presents an automated method for creating a first-draft thesaurus from raw text. It describes natural processing steps of tokenization, surface syntactic analysis, and syntactic attribute extraction. From these attributes, word and term similarity is calculated and a thesaurus is created showing important common terms and their relation to each other, common verb--noun pairings, common expressions, and word family members.
The techniques are tested on twenty different corpora ranging from baseball newsgroups, assassination archives, medical X-ray reports, abstracts on AIDS, to encyclopedia articles on animals, even on the text of the book itself. The corpora range from 40,000 to 6 million characters of text, and results are presented for each in the Appendix.
The methods described in the book have undergone extensive evaluation. Their time and space complexity are shown to be modest. The results are shown to converge to a stable state as the corpus grows. The similarities calculated are compared to those produced by psychological testing. A method of evaluation using Artificial Synonyms is tested. Gold Standards evaluation show that techniques significantly outperform non-linguistic-based techniques for the most important words in corpora.
Explorations in Automatic Thesaurus Discovery includes applications to the fields of information retrieval using established testbeds, existing thesaural enrichment, semantic analysis. Also included are applications showing how to create, implement, and test a first-draft thesaurus.

โœฆ Table of Contents


Front Matter....Pages i-xiii
Introduction....Pages 1-5
Semantic Extraction....Pages 7-32
Sextant....Pages 33-68
Evaluation....Pages 69-100
Applications....Pages 101-135
Conclusion....Pages 137-148
Back Matter....Pages 149-305

โœฆ Subjects


Artificial Intelligence (incl. Robotics); Language Translation and Linguistics


๐Ÿ“œ SIMILAR VOLUMES


Automated Taxonomy Discovery and Explora
โœ Jiaming Shen, Jiawei Han ๐Ÿ“‚ Library ๐Ÿ“… 2022 ๐Ÿ› Springer ๐ŸŒ English

<span>This book provides a principled data-driven framework that progressively constructs, enriches, and applies taxonomies without leveraging massive human annotated data. Traditionally, people construct domain-specific taxonomies by extensive manual curations, which is time-consuming and costly. I