Book Details
Format
eBook
Pages
231
Language
English
Published
Jan 11, 2010
Publisher
John Benjamins Publishing Co
ISBN-10
1282154567
ISBN-13
9781282154568
Description
This text covers the emerging technologies of document retrieval, information extraction, and text categorization in a way which highlights commonalities in terms of both general principles and practical issues. It seeks to satisfy a need on the part of technology practitioners in the internet space, faced with having to make difficult decisions as to what research has been done and what the best practices are. It is not intended as a vendor guide (such things are quickly out of date), or as a recipe for building applications (such recipes are very context-dependent). But it does identify the key technologies, the issues involved, and the strengths and weaknesses of the various approaches. There is also a strong emphasis on evaluation in every chapter, both in terms of methodology (how to evaluate) and what controlled experimentation and industrial experience have to tell us.
Genres
Science & Technology