Explorations in Automatic Thesaurus Discovery

Springer Science & Business Media, Jul 31, 1994 - Computers - 305 pages

Explorations in Automatic Thesaurus Discovery presents an automated method for creating a first-draft thesaurus from raw text. It describes natural processing steps of tokenization, surface syntactic analysis, and syntactic attribute extraction. From these attributes, word and term similarity is calculated and a thesaurus is created showing important common terms and their relation to each other, common verb--noun pairings, common expressions, and word family members.
The techniques are tested on twenty different corpora ranging from baseball newsgroups, assassination archives, medical X-ray reports, abstracts on AIDS, to encyclopedia articles on animals, even on the text of the book itself. The corpora range from 40,000 to 6 million characters of text, and results are presented for each in the Appendix.
The methods described in the book have undergone extensive evaluation. Their time and space complexity are shown to be modest. The results are shown to converge to a stable state as the corpus grows. The similarities calculated are compared to those produced by psychological testing. A method of evaluation using Artificial Synonyms is tested. Gold Standards evaluation show that techniques significantly outperform non-linguistic-based techniques for the most important words in corpora.
Explorations in Automatic Thesaurus Discovery includes applications to the fields of information retrieval using established testbeds, existing thesaural enrichment, semantic analysis. Also included are applications showing how to create, implement, and test a first-draft thesaurus.

Preview this book »

INTRODUCTION	1

SEMANTIC EXTRACTION	7

SEXTANT	33

EVALUATION	69

APPLICATIONS	101

CONCLUSION	137

PREPROCESSORS	149

SEMANTIC CLUSTERING	163

AUTOMATIC THESAURUS GENERATION	171

CORPORA TREATED	181

INDEX	294

Copyright

Bibliographic information

Title	Explorations in Automatic Thesaurus Discovery Volume 278 of The Springer International Series in Engineering and Computer Science
Author	Gregory Grefenstette
Edition	illustrated
Publisher	Springer Science & Business Media, 1994
ISBN	0792394682, 9780792394686
Length	305 pages
Subjects	Computers › Software Development & Engineering › General Computers / Intelligence (AI) & Semantics Computers / Natural Language Processing Computers / Software Development & Engineering / General

Export Citation	BiBTeX EndNote RefMan

About Google Books - Privacy Policy - Terms of Service - Information for Publishers - Report an issue - Help - Google Home

Books

Explorations in Automatic Thesaurus Discovery

Contents

Other editions - View all

Common terms and phrases

References to this book

Bibliographic information