Part of Speech n-Grams and Information Retrieval

Research output: Contribution to journalJournal articleResearchpeer-review

Efforts to use linguistics in information retrieval (IR) were initiated in the 1980s, and intensified in the 1990s, reporting performance benefits (see the overviews by Smeaton 1986 & 1999, Karlgren 1993, and Tait 2005). After that time, these efforts decreased: baseline system performance improved, and the cost associated with linguistic processing was not worth the small benefits over the already improved baselines (Tait, 2005). At present, most research on linguistics for IR tends to be geared towards domain-specific IR applications that seem to benefit more from linguistics, like question-answering (Tait & Oakes 2006). Although such applications are important, they should not limit the scope of research into linguistics for IR. In this work, we present an alternative use of linguistics, part of speech information in particular, to compute a term weight of informative content. This term weight is a novel application of linguistics to IR, and can benefit retrieval performance of general IR systems.
Original languageEnglish
JournalRevue Francaise de Linguistique Appliquee
VolumeXIII
Issue number2008/1
Pages (from-to)9-22
ISSN1386-1204
Publication statusPublished - 2008
Externally publishedYes

ID: 38240584